mrkeyoor.com_
Tue 06 Oct 06:35 UTC
Automationevaluationupdated 06 Oct 2026

Jev-cu review

Jev-cu's main documentation is in Chinese, and there is no English README. It is a Codex Computer Use skill that sends text from a desktop accessibility tree to TypeSafe's Jev model, asks it to choose the next control and action, then lets local policy code decide whether Codex may execute the step. Screenshots stay local, which reduces what leaves the machine but rules out work that depends on visual layout.

Verdict

Our Jev-cu run installed 0 packages and passed all 22 tests in 3 seconds, but the useful path still depends on Codex Desktop and a TypeSafe key. Use it for a controlled macOS experiment where the target controls have clear text labels and a person can approve sensitive steps. Keep it out of unattended or visually driven workflows until the stale-index and response-body timeout paths are closed.

We ran it

Lab card: what happened when we ran Jev-cuScreenshot of Jev-cu (github.com/Sac-Y/Jev-cu)
Install✓ · 5s0 packages · 1 MB
Buildn/ano build script
Tests✓ · 3s22 passed · 0 failed of 22 (node:test)
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo19 files~1,179 lines of source · 0.1 MB · 0 CI workflows · tests dir

Answers from our run

Does Jev-cu build from source?

Dependencies installed in 5 seconds (0 packages), and the project has no separate build step. We cloned commit 52d32ac into a clean Debian container with 3 CPUs and no project-specific setup.

Do Jev-cu's tests pass?

Yes: 22 of 22 passed when we ran the project's own test command (node:test). Some failures need services or credentials a bare container does not have.

Does Jev-cu have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use Jev-cu?

Teams that require English onboarding: the README, skill guide, runtime notes, and most code comments are in Chinese, with no English README.

What are the alternatives to Jev-cu?

Browser Use, OmniParser, Cua. Our Jev-cu run installed 0 packages and passed all 22 tests in 3 seconds, but the useful path still depends on Codex Desktop and a TypeSafe key.

Setup3/55-second install; live use needs Codex Desktop and a TypeSafe key
Docs3/5Clear Chinese safety guide, but no English README
Community2/5616 stars and nine open issues and PRs after six commits
Maturity2/5v0.1.0, no release or CI, and two open runtime faults

Who it’s for

Codex Desktop users who need help choosing among clearly labeled controls in a macOS app.
Teams testing text-only computer use and willing to supply a TypeSafe API key.
Developers who want dry-run previews, an app allowlist, and local confirmation rules around model-selected actions.
Researchers who can extend a small JavaScript loop and evaluate it against their own accessibility snapshots.

Who it’s NOT for

Teams that require English onboarding: the README, skill guide, runtime notes, and most code comments are in Chinese, with no English README.
Windows, Linux, or browser automation teams seeking a supported path: the published skill covers macOS App accessibility data and calls browser work a local experiment.
Work driven by canvas layout, 3D scenes, icons, or visual quality: the skill explicitly says those jobs do not fit its text-only approach.
Unattended workflows that must delete, send, pay, authorize, upload, install, or change system settings: those labels stop at a human confirmation gate.
Reliability-sensitive automation that cannot tolerate an index changing during a model call or a stalled response body: pull request 3 and issue 12 document those open paths.

Setup reality

Our sandbox installed commit 52d32ac in 5 seconds, added 0 packages, and used 1 MB on disk. There was no build script, so that step was skipped. Node's test runner finished in 3 seconds with 22 passed and 0 failed; npm audit found 0 known vulnerabilities.

Live use needs Codex Desktop's cua_repl, a TYPESAFE_API_KEY, and the skill copied into the local Codex skills directory. The API-backed P0 evaluation also needs that key. The guide says English goals work best even though the documentation is Chinese.

The runtime guidance covers macOS App accessibility data. It starts in dry-run mode, limits candidates to 40 by default, and asks callers to provide a result check. The repository has no CI workflow or Dockerfile, so the 22 local tests are the visible automated gate.

Jev chooses one labeled control while Codex holds the mouse

Jev-cu inserts a narrow decision step into Codex Computer Use. Codex reads the current accessibility tree, the text representation that desktop apps expose to assistive tools, and hands Jev a ranked list of controls. Jev answers four questions: which control to target, which action to take, whether the goal looks complete, and whether the next step looks risky. Codex still owns observation and execution. The remote model never receives a screenshot.

That split keeps the model inside a small box. The default candidate list tops out at 40 controls, labels sent to Jev are clipped to 120 characters, and recent actions are limited to six entries. Codex supplies any text, keyboard key, scroll direction, or coordinates needed by the chosen action.

A 0.20 risk score stops the action for confirmation

A Jev risk probability of 0.20 or higher returns confirm instead of executing. Local code also stops on labels associated with deletion, sending, payment, authorization, sharing, installation, or system settings. Apps outside the seven-name default allowlist take the same path. Confidence below 0.30 stops the loop; middling confidence escalates, with a slightly lower threshold for Calendar, Calculator, TextEdit, and Figma.

Dry-run is on by default and previews only one step. It does not pretend that later windows will match the first observation. A caller may add a verify function that checks the full accessibility state before the next step and after the final action. That is the strongest part of the design: the model saying done cannot overrule a failed result check. The guide also warns that an app allowlist is not blanket permission for every write inside that app.

What happened when we ran it

Our 3-CPU, 8 GB sandbox installed commit 52d32ac in 5 seconds. npm added 0 packages, and the installed checkout occupied 1 MB. The repository has 19 files and about 1,179 lines of source, so the result matches what the package file suggests: this is a small private JavaScript package with no dependency tree to pull in. npm audit reported 0 known vulnerabilities.

There is no build script or target, so our build step was skipped rather than passed. node:test completed in 3 seconds with 22 passed and 0 failed of 22. Those tests cover parsing, candidate ranking, policy decisions, Chinese accessibility roles, dry-run behavior, and simulated execution checks. The sandbox had no secrets, so this result does not cover the TypeSafe API or the API-backed P0 evaluation.

The 5-second install is the easy part

The install script copies the skill into ~/.codex/skills/jev-cu, and a new Codex session must load it. A live loop runs inside Codex Desktop's cua_repl, imports the repository module, binds to a named app, and reads that app's accessibility state. Jev needs TYPESAFE_API_KEY in the environment or .env.local.

The examples use English goals because the author says Jev is most accurate with them. The published runtime guide starts with Calendar, a dry run, and at most 2 steps. It also warns that the REPL normally times out after 30 seconds, while an API call plus desktop observations may need 60 seconds. If the outer call times out, the operator must inspect the app and trace before retrying because an action may already have happened.

Text-only input reduces exposure and loses visual evidence

Jev receives candidate text plus at most 1,500 characters of context, rather than a screenshot or the whole desktop state. It is not redaction. The guide explicitly warns that stripping URLs and shortening labels can leave private text in the payload, so callers still have to choose which controls and state lines leave the machine.

The same boundary creates blind spots. The skill rules out canvas layout, 3D modeling, and visual-quality judgments. Its parser recognizes a fixed collection of accessibility roles and currently normalizes five Chinese role names. An icon-only control, an unusual localized role, or a custom canvas may never enter the candidate set. OmniParser is a better comparison for those screens because it starts with pixels; Jev-cu starts with labels.

Two open execution faults rule out unattended use

Issue 12 shows that the 60,000 ms request timer is cleared when response headers arrive, before the JSON body has finished loading. A server that starts a response and then stalls can therefore hold the loop beyond its configured timeout. The current source still parses the body after clearing that timer. Pull request 13 proposes keeping the abort window active through body parsing, but it remains open.

Pull request 3 covers a more serious desktop race. Jev-cu observes the interface, waits for the remote decision, and then executes the returned accessibility index without a fresh observation immediately before the action. If the interface changes during that wait, the index can refer to a different control. The proposed fix rechecks the full state, at the cost of another observation and conservative stops. That protection is not in commit 52d32ac.

Six commits and no release keep v0.1.0 in trial territory

The repository was created on September 18, 2026, and its last push was September 22. GitHub showed 616 stars and 9 open issues and pull requests on October 6, split into four issues and five pull requests. There is no tagged release. The repo also has no CI workflow or Dockerfile, although its checked-in test directory gave us 22 passing tests. Activity arrived quickly, including fixes for Windows line endings, Chinese role parsing, and full-label safety checks.

That is enough evidence for a focused experiment, not for background automation that can change user data. Jev-cu has a sensible control boundary and unusually candid safety notes, yet the current loop still has a stale-index window and an incomplete network timeout. Keep dryRun enabled, supply an explicit result check, and require a person at every sensitive gate. If that operating model sounds too restrictive, this version is the wrong tool.

Alternatives

ProjectWhat it isPick it when
Browser Use gh↗A Python framework for agents that operate websites through a browser.pick this instead when the entire job is on the web and you want a browser-focused agent rather than a Codex Desktop skill.
OmniParserA vision-based screen parser that turns screenshots into elements a GUI agent can use.pick this instead when icons, canvas content, or screen position matter more than keeping screenshots out of the model input.
Cua gh↗A computer-use stack for cross-OS drivers, isolated environments, and evaluation fleets.pick this instead when you need computer-use infrastructure across machines rather than one small decision loop inside Codex Desktop.

What people are saying

  1. [velocity-scout] Sac-Y/Jev-cu

Sources

  1. Jev-cu repository
  2. Jev-cu README
  3. Jev-cu skill guide
  4. Jev-cu runtime guide
  5. Jev-cu policy source
  6. Response-body timeout issue 12
  7. Stale-index revalidation pull request 3
  8. Browser Use repository

More automation reviews

jev-browser-use · herdr-projects · repopilot · mobile-jev · wx_channels_download · jianying-headless · the whole board →