mrkeyoor.com_
Wed 07 Oct 07:31 UTC
LLM Toolsevaluationupdated 07 Oct 2026

jev-skill review

Jev Skill is an English-first, Chinese-translated collection of coding-agent skills for classification, scoring, evidence review, and choosing among bounded actions with Jev. It gives Codex, Claude Code, and OpenCode reusable instructions and JSON templates while leaving evidence gathering and execution to the host agent.

Verdict

Our run built in 4 seconds and passed 170 tests, but pytest still exited 1 on a duplicate-module collection error, so Jev Skill is usable material with a packaging wrinkle rather than a clean release gate. Pick it when you want detailed Jev workflows and will pin either the 11-skill v0.2.0 release or the five-skill source preview deliberately. Avoid it if your team needs one stable install surface or might mistake a simulated judgment for a real Jev call.

We ran it

Lab card: what happened when we ran jev-skillScreenshot of jev-skill (github.com/wuyoscar/jev-skill)
Install✓ · 26s36 packages · 38 MB
Build✓ · 4s
Tests✗ · 19s170 passed · 0 failed · 1 errors of 171 (pytest)
Known vulns0(pip-audit)
Repo489 files~5,838 lines of source · 18 MB · 2 CI workflows · tests dir

Answers from our run

Does jev-skill build from source?

Dependencies installed in 26 seconds (36 packages), and the build succeeded in 4 seconds. We cloned commit 01bd940 into a clean Debian container with 3 CPUs and no project-specific setup.

Do jev-skill's tests pass?

Yes: 170 of 171 passed when we ran the project's own test command (pytest), with 1 collection error. Some failures need services or credentials a bare container does not have.

Does jev-skill have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use jev-skill?

Teams expecting the default branch and latest release to have the same shape: main is a five-skill preview, while v0.2.0 contains eleven entry points.

What are the alternatives to jev-skill?

Jev MCP, SemDecide, Awesome Jev by TypeSafe. Our run built in 4 seconds and passed 170 tests, but pytest still exited 1 on a duplicate-module collection error, so Jev Skill is usable material with a packaging wrinkle rather than a clean release gate.

Setup3/5Small install, but source choice and host folder copying need care
Docs5/5Detailed bilingual install, safety, migration, and workflow guides
Community4/5579 stars and recent merged PRs, with no open issues or PRs
Maturity3/5CI and a tagged release exist, but main has an unreleased layout

Who it’s for

Codex, Claude Code, or OpenCode users who want reusable Jev decision workflows inside a project.
Teams that need explicit review paths, source-linked judgments, and dry-run request validation.
Researchers comparing real Jev calls with clearly labeled agent or model simulation.
Skill authors who want worked templates for triage, document review, evaluation, and bounded actions.

Who it’s NOT for

Teams expecting the default branch and latest release to have the same shape: main is a five-skill preview, while v0.2.0 contains eleven entry points.
Users who want real Jev results without a provider account: API mode needs OpenRouter or TypeSafe credentials and paid network calls; simulation is labeled as simulation and carries no Jev confidence.
Operators who treat a model judgment as permission to act: the skill repeatedly assigns execution, authorization, and final verification to the host.
Release gates that require one clean repository-wide pytest run: our run had 170 passing tests but exited 1 during collection because two experiment files share the test_pilot.py basename.
Anyone expecting a CLI install to make skills appear automatically: the wheel installs jev-decide, while host skill folders must be copied and discovered separately.

Setup reality

Our sandbox installed commit 01bd940 in 26 seconds, adding 36 packages and using 38 MB. The build succeeded in 4 seconds. Pytest exited 1 after 19 seconds: 170 tests passed, none failed, 272 subtests passed, and 1 collection/setup error prevented a clean 171-test run. Pip-audit found 0 known vulnerabilities.

Before copying anything, choose between the tagged v0.2.0 release with 11 skills and the five-skill preview on main. The documented default is project-local installation for Codex, Claude Code, or OpenCode. Real API mode needs Python 3.10+, an existing uv or pipx, and either an OpenRouter or TypeSafe key.

Simulation mode needs neither the CLI nor a Jev key, but it must label the result and leave probability and confidence null. Offline dry runs validate request shape only. They do not prove credentials, paid access, model quality, or native skill invocation inside the host.

Main has five skills, while v0.2.0 has eleven

Jev Skill's first setup decision is which product you mean. The default branch presents five folders: jev, jev-triage, jev-documents, jev-eval, and jev-act. Its install guide says this arrangement is a source preview. The latest tagged release, v0.2.0, retains eleven entry points. The current pyproject.toml reports package version 0.2.1, but that number does not turn the preview into a published release. Pin a tag or exact commit before copying files.

That split is easy to miss because the README is also a large catalog. It lists projects, scenarios, reference material, and recorded examples alongside the installable skills. The collection can teach workflow design even if you never install its CLI. If you do install it, the source choice affects skill names, migration, and what your host discovers. A team that follows main without review may receive a different interface from a teammate using v0.2.0.

The host agent still gathers evidence and takes action

The five preview skills divide work by judgment type. The general jev skill covers custom choices, routing, and checkpoints. jev-triage handles records, jev-documents maps claims to source evidence, jev-eval reviews outputs, and jev-act selects among legal actions. In each case Jev receives bounded candidates or rubric questions. The surrounding coding agent collects context, decides what may run, and verifies the result.

That boundary is well written and worth preserving. A selection is not authorization, confidence is not accuracy, and an evaluation lead is not merge approval. The skills require an unknown or review path when evidence is missing. They also tell the host to separate trusted rules from untrusted text and to wait for new observations before dependent judgments. Those instructions make the collection safer than a loose prompt pack, though enforcement still depends on the host following them.

What happened when we ran it

Our sandbox installed commit 01bd940 in 26 seconds. The Python environment gained 36 packages and used 38 MB on disk. The package build completed successfully in 4 seconds, and pip-audit reported 0 known vulnerabilities. The checkout contained 489 files and about 5,838 lines of source, with a tests directory and two GitHub Actions workflow files.

Pytest ran for 19 seconds and ended with exit code 1. It reported 170 passing tests, no failed assertions, 272 passing subtests, and 1 collection/setup error in the 171-test run. The error is specific: pytest imported docs/experiments/context-pilot/test_pilot.py as module test_pilot, then encountered docs/experiments/public-pr-pilot/test_pilot.py under the same module name. Collection stopped on that file mismatch.

This is a narrower problem than a failing behavior test, yet it still matters. A clean checkout using the documented broad test command did not finish successfully in our Debian container. The log itself suggests unique basenames or clearing cached bytecode as avenues, but it does not prove which repository change the maintainers intend. The fair result is 170 passing tests plus a collection defect, not a green suite and not an assertion failure.

Installation begins with a host and source choice

The 26-second package install does not place the skills into your agent automatically. The wheel supplies the jev-decide command. Codex looks under .agents/skills/, Claude Code under .claude/skills/, and OpenCode under .opencode/skills/. The guide recommends project-local installation by default, complete-folder copies, conflict checks before replacement, and an exact source commit in the final report.

Real API mode adds Python 3.10 or newer plus an existing uv or pipx. The installer is told not to use sudo, modify system Python, or rewrite shell profiles. It checks only whether OPENROUTER_API_KEY or TYPESAFE_API_KEY exists and must never print the value. Offline --dry-run calls confirm that JSON requests can be parsed and mapped. They do not authenticate an account or spend money.

The five-skill copy route itself needs no Node/npm, Vercel account, gateway, or MCP server. Verification is more involved than listing folders: the guide checks each copied SKILL.md, local links, selected example files, merged-mode requests, and host discovery. Even then, file layout does not prove a native slash command or trigger works. The documentation makes that limitation explicit, which is more useful than declaring success after a copy.

Simulation is a separate mode, not free Jev

Without a provider key, the collection offers simulation by the current agent or another model the user explicitly selects. A simulated result must say agent_simulation or model_simulation, set jev_called to false, and leave probability and confidence null. That makes a no-key trial useful for designing the workflow without pretending a different model came from Jev.

Real mode sends supplied evidence to either OpenRouter or TypeSafe and may cost money. The user chooses the provider once, and an error does not authorize a silent switch. The skill asks before private data leaves the machine or a paid call occurs. This careful ceremony is appropriate for a reusable skill, but it adds steps for anyone hoping to paste one prompt and immediately receive a verified external judgment.

Two workflows and September pull requests show active maintenance

GitHub showed 579 stars, no open issues or pull requests, and a September 30, 2026 last push when checked on October 7. The repository has two workflow files, including its test workflow. Recent merged pull requests addressed Windows text handling, provider policy, contributor automation, and experiment documentation. Release v0.2.0 was published on September 21 under the MIT license.

The project is active, documented, and cheap to inspect locally. Its weak point is version clarity: the README asks an installing agent to explain a release-versus-preview distinction before doing any work, which is a burden the package should eventually remove. For now, pin the source, keep real and simulated output separate, and treat our pytest collection error as an unresolved release-quality blemish.

Alternatives

ProjectWhat it isPick it when
Jev MCPAn MCP server exposing named Jev tools for classification, reranking, review, and gates.pick this instead when your clients already speak MCP and you want callable tools rather than a large workflow library.
SemDecideA CLI for typed semantic decisions in shell pipelines and CI jobs.pick this instead when JSONL and Unix pipelines matter more than coding-agent skill discovery.
Awesome Jev by TypeSafeAn evidence-oriented catalog of Jev patterns, prompts, and starter examples.pick this instead when you want a lighter reading list rather than installable host-specific skills.

What people are saying

  1. [velocity-scout] wuyoscar/jev-skill

Sources

  1. Jev Skill README at commit 01bd940
  2. Agent installation guide at commit 01bd940
  3. Core Jev skill instructions at commit 01bd940
  4. Jev Skill v0.2.0 release
  5. Jev Skill GitHub repository

More llm tools reviews

minorun-marp-skill · jev-pruner · llm-d-router · Rapid-MLX · simple-jev · kev · the whole board →