jev-align turns labels into a versioned Jev decision function
jev-align starts with a typed task and a local dataset, then searches for examples where the current Jev definition is uncertain. A round mixes ambiguous rows with an audit sample, asks you to confirm labels and optional rationales, and passes feedback into GEPA. GEPA proposes definition changes, while the CLI shows score and certainty movement before you accept, reject, rewind, or pause. Binary, multiclass, multilabel, and ordered-score tasks all use the same guided loop.
The default pool uses the first 1,000 rows or the whole file when it is smaller. A round can request 5, 10, 15, or 20 annotations, reserve an optional 20% holdout, and give GEPA a default budget of 300 metric calls. Those controls make the process legible, but they do not turn labels into ground truth automatically. The README tells users to review every label, including synthetic or agent-assisted ones, and no higher training score accepts a proposal without a person.
The 866 MB environment is heavier than the small CLI suggests
Our measured checkout had 82 files, about 18,470 lines of source, and occupied 10.8 MB. Installation completed in 29 seconds, pulling 160 packages and expanding the environment to 866 MB. That footprint comes before your datasets, saved runs, or provider-side usage. The repository includes 1 CI workflow file and a tests directory, but no Dockerfile. Python 3.11 or newer is the documented base, and the package currently identifies itself as alpha software.
The CLI accepts CSV, Parquet, and JSONL input. Jev evaluation can go through TypeSafe AI, Vercel AI Gateway, or Cloudflare Workers AI. GEPA's reflection model is separate and can use OpenAI, Anthropic, Gemini, a local vLLM endpoint, or another LiteLLM provider. A typical run therefore needs two model routes with different credentials, costs, data policies, and failure behavior. Selecting a gateway for Jev does not also settle where reflection examples and rationales go.
What happened when we ran it
Our sandbox installed commit 520aec3 in 29 seconds, built it in 1 second, and completed the test step successfully in 13 seconds. The container had 3 CPUs, 8 GB of RAM, Python 3.12 on Debian, no secrets, and no elevated privileges. Pip-audit reported 2 known vulnerabilities across the 160 installed packages. Our measurement method did not contact Jev, a reflection model, or the public function registry.
The successful suite is evidence that the repository's supplied checks worked in our environment. It does not measure label quality, optimization gains, provider latency, or decision accuracy. The 2 audit findings are also only a count from the supplied run; the measurement block does not name packages or advisories. A production assessment should identify those dependencies, determine whether the vulnerable paths are reachable, and rerun the audit against the exact locked environment you plan to deploy.
Issue #3 makes the displayed holdout score unsafe to trust blindly
The only open issue reports that a held-out batch containing one true negative can display F1 0.000 even when the prediction is correct. According to the report, the metric treats undefined precision and recall as zero, then stores and displays the resulting score. That does not prove every task is mis-scored, but it is enough to require a known dataset containing positive and negative cases before you use the dashboard to compare definitions.
The same issue reports that TypeSafe's SDK path has retry policy while the Vercel and Cloudflare gateway code does not. A 429 or server error can abort a round, and the reported message points users toward credentials and model availability even for rate limiting. The issue also describes a round consuming 40 metric calls without producing a reflection candidate. These are user-reported findings rather than results from our sandbox, but each maps to a concrete acceptance test.
Public functions can include labels and rationales
Saved functions live under .jev-align/runs/ and include the accepted definition, input signature, backend, learning configuration, labels, and split information. Publishing to ai-functions.dev can also include labeled inputs and optional rationales. The README says it excludes unlabeled source rows, local paths, credentials, and unaccepted proposals. The registry is public, and its annotations are not guaranteed to have been created or reviewed by a person.
Unpublishing stops discovery and future public pulls, but it cannot revoke copies already downloaded. That makes the review screen a data-release boundary, not merely a sharing dialog. Release v0.1.6 and the latest push both landed on October 3, 2026. GitHub showed 304 stars and 1 open issue. jev-align is active and unusually candid about its dangerous label-free squash mode. The right trial uses non-sensitive data, fixed evaluation examples, and explicit checks for the metric and retry failures already reported.

