mrkeyoor.com_
Wed 07 Oct 06:44 UTC
AI Toolsevaluationupdated 07 Oct 2026

openJev-verdict-2.0 review

openJev-verdict-2.0 is a Python research project that turns a block of text into bounded choices, scores, or yes/no-style decisions with probabilities instead of generating prose. The repository covers two related models, a 151M-parameter Verdict checkpoint and a separate Verdict 2.0 architecture for typed software workflows.

Verdict

Our run installed 98 packages and used 5,719 MB, but 3 tests failed, so openJev-verdict-2.0 is easier to inspect than to trust as a dependency. Its receipts, typed schemas, and evaluation code make it useful to researchers who will validate the model path themselves. Application teams should wait for a retrievable Verdict 2.0 checkpoint, a green suite, and an unambiguous license file.

We ran it

Lab card: what happened when we ran openJev-verdict-2.0Screenshot of openJev-verdict-2.0 (heman10x-ngu.github.io/openJev-verdict-2.0)
Install✓ · 82s98 packages · 5719 MB
Build✓ · 5s
Tests✗ · 75s30 passed · 3 failed · 9 skipped of 33 (pytest)
Known vulns0(pip-audit)
Repo189 files~14,813 lines of source · 50.8 MB · 0 CI workflows · tests dir

Answers from our run

Does openJev-verdict-2.0 build from source?

Dependencies installed in 82 seconds (98 packages), and the build succeeded in 5 seconds. We cloned commit bff2856 into a clean Debian container with 3 CPUs and no project-specific setup.

Do openJev-verdict-2.0's tests pass?

Not all of them: 30 of 33 passed and 3 failed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does openJev-verdict-2.0 have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use openJev-verdict-2.0?

Teams requiring a green release gate: 3 tests failed in our commit bff2856 run, and the repository has 0 CI workflow files.

What are the alternatives to openJev-verdict-2.0?

Laya, jevos, Ollaya. Our run installed 98 packages and used 5,719 MB, but 3 tests failed, so openJev-verdict-2.

Setup2/5Build passed, but 5,719 MB and 3 failures raise the cost
Docs3/5Detailed runbook and receipts, but the docs freshness test fails
Community3/5293 stars and 44 forks, with 3 open issues after the last push
Maturity1/5No release or CI, a failed suite, and disputed artifact access

Who it’s for

ML engineers studying typed query shapes across a 189-file research repository for routing, triage, or policy checks.
Researchers who want prediction receipts, evaluation scripts, calibration data, and failure examples in the same repository.
Browser-inference experimenters willing to inspect the included WebGPU demo and manage model artifacts themselves.
Teams prepared to repair packaging and test drift before putting the engine behind an application.

Who it’s NOT for

Teams requiring a green release gate: 3 tests failed in our commit bff2856 run, and the repository has 0 CI workflow files.
Developers who need the Verdict 2.0 checkpoint to arrive through the documented clone path: open issues 2 and 4 report that the 598 MB Git LFS object returns 404.
Organizations that need an SPDX-recognized Apache-2.0 grant without further review: GitHub reports NOASSERTION, while open issue 2 says the shortened LICENSE omits conditions from the standard text.
Buyers expecting a versioned, drop-in package: GitHub has no published release, the package identifies itself as rlcd 0.1.0, and the README quickstart stops at tokenizer setup.

Setup reality

Our fresh Debian sandbox installed commit bff2856 in 82 seconds, adding 98 packages and using 5,719 MB on disk. The build succeeded in 5 seconds. Pytest then exited 1 after 75 seconds: 30 passed, 3 failed, and 9 were skipped. Pip-audit found 0 known vulnerabilities.

Local inference needs Python 3.10 or newer plus PyTorch, Transformers, GLiClass, ONNX tooling, and model artifacts. The README points Verdict users to a Hugging Face checkpoint, while Verdict 2.0 uses a Git LFS pointer in the repository. Two open issues report that the 598 MB LFS object cannot be downloaded.

The training runbook targets an NVIDIA GTX 1660 Ti and documents different flags for base and large models. The browser demo expects Chrome or Edge with WebGPU. There is no Dockerfile or CI workflow, so you must supply the container and automated checks yourself.

Three typed queries replace free-form generation

The repository's 189 files define 3 query shapes: Choice selects one option, Score assigns an ordered level, and Noul returns true, false, or insufficient evidence. That is a useful boundary for software that must choose a queue, rate severity, or decide whether a claim has support. The model supplies probabilities, while application code keeps control of thresholds and actions. No parser has to recover a decision from a paragraph.

Each choice or score accepts 2 to 24 candidates, and the result records the selected item, its probability, calibration status, and latency. The engine batches several questions into one forward call. It can use PyTorch or a local ONNX file, and it adds an explicit insufficient-evidence outcome. These details make the repository more interesting than a leaderboard image because you can inspect how a decision reaches ordinary Python code.

Verdict 1.4 and Verdict 2.0 are different deliverables

The README describes 2 models under one project name. Verdict is the 151M GLiClass-based checkpoint hosted as heman10x/rlcd-modernbert-151m. Verdict 2.0 is a separate typed-workflow architecture evaluated by the verdict2 code. Its weights are represented by artifacts/verdict2-base/model.pt, while the downloadable artifact manifest points to the earlier GLiClass model. A developer can easily follow one path while believing it is the other.

Open issues 2 and 4 report that the Verdict 2.0 Git LFS object is missing, with a 404 for the 598 MB checkpoint. Issue 2 also says the similarly named Hugging Face path returned 401 while the Verdict checkpoint remained available. Those are user reports, not failures from our sandbox. They still matter because the repository's main claim concerns Verdict 2.0, and a pointer file cannot perform inference.

What happened when we ran it

Our sandbox installed commit bff2856 in 82 seconds. The environment pulled 98 packages and occupied 5,719 MB, far more than the 50.8 MB checkout. The build completed successfully in 5 seconds. Our measurement setup was a fresh Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit reported 0 known vulnerabilities in the installed environment.

Pytest failed with exit code 1 after 75 seconds. It reported 30 passed, 3 failed, and 9 skipped. That result is specific to commit bff2856 and the sandbox above. It does not measure the README's accuracy, calibration, or latency claims. It does show that a fresh installation and successful build do not produce a green repository test run.

The 3 failures reach documentation and real inference

Our run ended with 3 failed tests, including a documentation receipt check that said the checked-in prose had drifted and instructed the maintainer to run python scripts/render_receipts.py. The repository contains reports, prediction logs, and scripts meant to tie claims back to generated evidence. Readers cannot assume every displayed figure was rendered from the current receipts, even though 30 other tests passed.

The real GLiClass inference test expected confidence above 0.7 and received 0.47440735212575424. The formatting test expected a bare label such as Cancel active subscription but got It is Cancel active subscription. That prefix matches the README's description of a v1.4 inference change, yet the test still rejects it. The log establishes the mismatch. It does not establish whether the implementation or expectation should change.

Zero CI workflows leave those failures outside an automated gate

Our scan found 0 CI workflow files and no Dockerfile, although the 189-file repository does include a tests directory. GitHub showed 293 stars, 44 forks, and 3 open issues and pull requests on October 7, 2026. All 3 were issues. The latest commit was pushed on September 20, three days before the newest open issue, so the queue contains reports that arrived after the last code change.

GitHub also returned no latest release. The Python metadata calls the package rlcd version 0.1.0 rather than the repository name. Licensing needs attention too: the README says Apache 2.0, while GitHub identifies the included shortened text as NOASSERTION, and issue 2 disputes whether it preserves the standard terms. A company should resolve that conflict before redistributing the code or weights.

A 5,719 MB environment is steep for an early research repo

Python 3.10 or newer is required, and the default dependency list includes PyTorch, Transformers, GLiClass, ONNX, ONNX Runtime, Pydantic, NumPy, and Accelerate. Training has a separate requirements file and a runbook built around a 6 GB NVIDIA GPU. Browser use adds WebGPU and model-file handling. None of that is unreasonable for model research, but it is a lot to own when 3 tests fail and the named Verdict 2.0 weight is disputed.

The successful 5-second build is still useful evidence. The source can be installed, its schemas are concrete, and much of its suite works. Start here if you want to inspect calibration, abstention, permutation tests, or typed decision interfaces and are willing to verify every artifact. Keep it off a production decision path until the checkpoint is retrievable, the 3 failures are resolved, and the license matches the promise in the README.

Alternatives

ProjectWhat it isPick it when
Laya gh↗A packaged typed-decision engine with Python, TypeScript, server, and multilingual paths.pick this instead when you want a published package, broader language support, and maintained client integrations.
jevosA local CPU decision server distributed as a binary with a TypeSafe-compatible API.pick this instead when a small operational surface and CPU deployment matter more than studying Python training code.
OllayaA local daemon that pulls and serves several open decision-model families behind one API.pick this instead when you want model management, a stable server process, and the freedom to compare several checkpoints.

What people are saying

  1. [velocity-scout] Heman10x-NGU/openJev-verdict-2.0

Sources

  1. openJev-verdict-2.0 repository and README
  2. Verdict 2.0 checkpoint and license report
  3. Git LFS checkpoint download report
  4. JevBench v1.4 results issue
  5. Commit bff2856

More ai tools reviews

embodied-jev · underclass · minecraft-agent · laya-coreml · CometixCode · editaplot2026 · the whole board →