Three typed queries replace free-form generation
The repository's 189 files define 3 query shapes: Choice selects one option, Score assigns an ordered level, and Noul returns true, false, or insufficient evidence. That is a useful boundary for software that must choose a queue, rate severity, or decide whether a claim has support. The model supplies probabilities, while application code keeps control of thresholds and actions. No parser has to recover a decision from a paragraph.
Each choice or score accepts 2 to 24 candidates, and the result records the selected item, its probability, calibration status, and latency. The engine batches several questions into one forward call. It can use PyTorch or a local ONNX file, and it adds an explicit insufficient-evidence outcome. These details make the repository more interesting than a leaderboard image because you can inspect how a decision reaches ordinary Python code.
Verdict 1.4 and Verdict 2.0 are different deliverables
The README describes 2 models under one project name. Verdict is the 151M GLiClass-based checkpoint hosted as heman10x/rlcd-modernbert-151m. Verdict 2.0 is a separate typed-workflow architecture evaluated by the verdict2 code. Its weights are represented by artifacts/verdict2-base/model.pt, while the downloadable artifact manifest points to the earlier GLiClass model. A developer can easily follow one path while believing it is the other.
Open issues 2 and 4 report that the Verdict 2.0 Git LFS object is missing, with a 404 for the 598 MB checkpoint. Issue 2 also says the similarly named Hugging Face path returned 401 while the Verdict checkpoint remained available. Those are user reports, not failures from our sandbox. They still matter because the repository's main claim concerns Verdict 2.0, and a pointer file cannot perform inference.
What happened when we ran it
Our sandbox installed commit bff2856 in 82 seconds. The environment pulled 98 packages and occupied 5,719 MB, far more than the 50.8 MB checkout. The build completed successfully in 5 seconds. Our measurement setup was a fresh Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. Pip-audit reported 0 known vulnerabilities in the installed environment.
Pytest failed with exit code 1 after 75 seconds. It reported 30 passed, 3 failed, and 9 skipped. That result is specific to commit bff2856 and the sandbox above. It does not measure the README's accuracy, calibration, or latency claims. It does show that a fresh installation and successful build do not produce a green repository test run.
The 3 failures reach documentation and real inference
Our run ended with 3 failed tests, including a documentation receipt check that said the checked-in prose had drifted and instructed the maintainer to run python scripts/render_receipts.py. The repository contains reports, prediction logs, and scripts meant to tie claims back to generated evidence. Readers cannot assume every displayed figure was rendered from the current receipts, even though 30 other tests passed.
The real GLiClass inference test expected confidence above 0.7 and received 0.47440735212575424. The formatting test expected a bare label such as Cancel active subscription but got It is Cancel active subscription. That prefix matches the README's description of a v1.4 inference change, yet the test still rejects it. The log establishes the mismatch. It does not establish whether the implementation or expectation should change.
Zero CI workflows leave those failures outside an automated gate
Our scan found 0 CI workflow files and no Dockerfile, although the 189-file repository does include a tests directory. GitHub showed 293 stars, 44 forks, and 3 open issues and pull requests on October 7, 2026. All 3 were issues. The latest commit was pushed on September 20, three days before the newest open issue, so the queue contains reports that arrived after the last code change.
GitHub also returned no latest release. The Python metadata calls the package rlcd version 0.1.0 rather than the repository name. Licensing needs attention too: the README says Apache 2.0, while GitHub identifies the included shortened text as NOASSERTION, and issue 2 disputes whether it preserves the standard terms. A company should resolve that conflict before redistributing the code or weights.
A 5,719 MB environment is steep for an early research repo
Python 3.10 or newer is required, and the default dependency list includes PyTorch, Transformers, GLiClass, ONNX, ONNX Runtime, Pydantic, NumPy, and Accelerate. Training has a separate requirements file and a runbook built around a 6 GB NVIDIA GPU. Browser use adds WebGPU and model-file handling. None of that is unreasonable for model research, but it is a lot to own when 3 tests fail and the named Verdict 2.0 weight is disputed.
The successful 5-second build is still useful evidence. The source can be installed, its schemas are concrete, and much of its suite works. Start here if you want to inspect calibration, abstention, permutation tests, or typed decision interfaces and are willing to verify every artifact. Keep it off a production decision path until the checkpoint is retrievable, the 3 failures are resolved, and the license matches the promise in the README.

