mrkeyoor.com_
Wed 07 Oct 06:42 UTC
AI Toolsevaluationupdated 07 Oct 2026

jev-align review

jev-align is an experimental Python CLI for turning labeled examples into reusable Jev decision functions. It finds uncertain rows, asks a person to label them, and uses GEPA plus a separate reflection model to propose a better typed definition without accepting that proposal automatically.

Verdict

Our jev-align run installed 160 packages into 866 MB, passed its build and tests, and still returned 2 known vulnerabilities, so a trial is justified but a blind rollout is not. The human approval loop and portable artifacts are thoughtful. Fix or work around the held-out metric and gateway retry defects before using its scores to steer an important decision function.

We ran it

Lab card: what happened when we ran jev-alignScreenshot of jev-align (pypi.org/project/jev-align)
Install✓ · 29s160 packages · 866 MB
Build✓ · 1s
Tests✓ · 13sran, no count parsed
Known vulns2(pip-audit)
Repo82 files~18,470 lines of source · 10.8 MB · 1 CI workflows · tests dir

Answers from our run

Does jev-align build from source?

Dependencies installed in 29 seconds (160 packages), and the build succeeded in 1 seconds. We cloned commit 520aec3 into a clean Debian container with 3 CPUs and no project-specific setup.

Do jev-align's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does jev-align have known vulnerabilities in its dependencies?

pip-audit flagged 2 known advisories in the dependency tree at the time of our run.

Who should not use jev-align?

High-stakes evaluation teams until they examine issue #3: it reports a correct all-negative held-out round displayed as F1 0.000.

What are the alternatives to jev-align?

DSPy, Label Studio, Argilla. Our jev-align run installed 160 packages into 866 MB, passed its build and tests, and still returned 2 known vulnerabilities, so a trial is justified but a blind rollout is not.

Setup3/5Fast checks, but 160 packages and two providers add setup
Docs4/5Workflow, providers, publishing, and dangerous squash mode are clear
Community3/5304 stars, one detailed open issue, and an October 3 push
Maturity2/5Alpha 0.1.6 with open metric, retry, and dependency concerns

Who it’s for

Teams building binary, multiclass, multilabel, or rubric-scoring decisions with Jev.
Developers who have CSV, Parquet, or JSONL examples and can review labels round by round.
Evaluators who want accepted definitions, labels, splits, and rationales saved as portable artifacts.
Python users comfortable managing both a Jev provider and a reflection-model provider.

Who it’s NOT for

High-stakes evaluation teams until they examine issue #3: it reports a correct all-negative held-out round displayed as F1 0.000.
Vercel or Cloudflare users who require built-in transient-error recovery: the same open issue reports no retry on those gateway paths.
Users who cannot send selected rows, labels, and rationales to a reflection-model provider during GEPA optimization.
Anyone treating lower uncertainty as proof of correctness: the README warns that squash mode can become confidently wrong.
Teams unable to remediate dependency findings: our pip-audit reported 2 known vulnerabilities.

Setup reality

Our sandbox installed commit 520aec3 in 29 seconds, pulling 160 packages and using 866 MB. The build succeeded in 1 second, and the tests succeeded in 13 seconds. Pip-audit reported 2 known vulnerabilities.

Python 3.11 or newer is required. A normal optimization needs credentials for TypeSafe, Vercel AI Gateway, or Cloudflare Workers AI, plus a separately configured OpenAI, Anthropic, Gemini, or other LiteLLM reflection model.

The CLI reads CSV, Parquet, or JSONL data and stores runs under .jev-align/runs/. Publishing can expose accepted definitions, labeled inputs, labels, splits, and rationales in a public registry; unpublishing cannot revoke copies already downloaded.

jev-align turns labels into a versioned Jev decision function

jev-align starts with a typed task and a local dataset, then searches for examples where the current Jev definition is uncertain. A round mixes ambiguous rows with an audit sample, asks you to confirm labels and optional rationales, and passes feedback into GEPA. GEPA proposes definition changes, while the CLI shows score and certainty movement before you accept, reject, rewind, or pause. Binary, multiclass, multilabel, and ordered-score tasks all use the same guided loop.

The default pool uses the first 1,000 rows or the whole file when it is smaller. A round can request 5, 10, 15, or 20 annotations, reserve an optional 20% holdout, and give GEPA a default budget of 300 metric calls. Those controls make the process legible, but they do not turn labels into ground truth automatically. The README tells users to review every label, including synthetic or agent-assisted ones, and no higher training score accepts a proposal without a person.

The 866 MB environment is heavier than the small CLI suggests

Our measured checkout had 82 files, about 18,470 lines of source, and occupied 10.8 MB. Installation completed in 29 seconds, pulling 160 packages and expanding the environment to 866 MB. That footprint comes before your datasets, saved runs, or provider-side usage. The repository includes 1 CI workflow file and a tests directory, but no Dockerfile. Python 3.11 or newer is the documented base, and the package currently identifies itself as alpha software.

The CLI accepts CSV, Parquet, and JSONL input. Jev evaluation can go through TypeSafe AI, Vercel AI Gateway, or Cloudflare Workers AI. GEPA's reflection model is separate and can use OpenAI, Anthropic, Gemini, a local vLLM endpoint, or another LiteLLM provider. A typical run therefore needs two model routes with different credentials, costs, data policies, and failure behavior. Selecting a gateway for Jev does not also settle where reflection examples and rationales go.

What happened when we ran it

Our sandbox installed commit 520aec3 in 29 seconds, built it in 1 second, and completed the test step successfully in 13 seconds. The container had 3 CPUs, 8 GB of RAM, Python 3.12 on Debian, no secrets, and no elevated privileges. Pip-audit reported 2 known vulnerabilities across the 160 installed packages. Our measurement method did not contact Jev, a reflection model, or the public function registry.

The successful suite is evidence that the repository's supplied checks worked in our environment. It does not measure label quality, optimization gains, provider latency, or decision accuracy. The 2 audit findings are also only a count from the supplied run; the measurement block does not name packages or advisories. A production assessment should identify those dependencies, determine whether the vulnerable paths are reachable, and rerun the audit against the exact locked environment you plan to deploy.

Issue #3 makes the displayed holdout score unsafe to trust blindly

The only open issue reports that a held-out batch containing one true negative can display F1 0.000 even when the prediction is correct. According to the report, the metric treats undefined precision and recall as zero, then stores and displays the resulting score. That does not prove every task is mis-scored, but it is enough to require a known dataset containing positive and negative cases before you use the dashboard to compare definitions.

The same issue reports that TypeSafe's SDK path has retry policy while the Vercel and Cloudflare gateway code does not. A 429 or server error can abort a round, and the reported message points users toward credentials and model availability even for rate limiting. The issue also describes a round consuming 40 metric calls without producing a reflection candidate. These are user-reported findings rather than results from our sandbox, but each maps to a concrete acceptance test.

Public functions can include labels and rationales

Saved functions live under .jev-align/runs/ and include the accepted definition, input signature, backend, learning configuration, labels, and split information. Publishing to ai-functions.dev can also include labeled inputs and optional rationales. The README says it excludes unlabeled source rows, local paths, credentials, and unaccepted proposals. The registry is public, and its annotations are not guaranteed to have been created or reviewed by a person.

Unpublishing stops discovery and future public pulls, but it cannot revoke copies already downloaded. That makes the review screen a data-release boundary, not merely a sharing dialog. Release v0.1.6 and the latest push both landed on October 3, 2026. GitHub showed 304 stars and 1 open issue. jev-align is active and unusually candid about its dangerous label-free squash mode. The right trial uses non-sensitive data, fixed evaluation examples, and explicit checks for the metric and retry failures already reported.

Alternatives

ProjectWhat it isPick it when
DSPy gh↗A framework for programming and optimizing language-model pipelines against metrics.pick this instead when you are optimizing multi-step generative programs rather than Jev typed decisions.
Label Studio gh↗A general data-labeling platform with configurable annotation interfaces.pick this instead when annotation operations and reviewer workflows matter more than automatic definition optimization.
ArgillaA collaboration platform for building and reviewing AI feedback datasets.pick this instead when you need shared dataset curation across a team before choosing a model or optimizer.

What people are saying

  1. [velocity-scout] sutro-sh/jev-align

Sources

  1. jev-align README
  2. Held-out F1 and gateway retry issue
  3. jev-align 0.1.6 release
  4. jev-align package metadata

More ai tools reviews

embodied-jev · underclass · minecraft-agent · laya-coreml · CometixCode · openJev-verdict-2.0 · the whole board →