Typed probabilities replace free-form answer parsing
OpenJev turns a state and a set of questions into fixed-shape answers. A question can be yes or no, a named choice, or a score with up to 10 levels. The server reads probabilities from designated answer tokens rather than asking a language model to write JSON. That removes one familiar failure mode: the answer cannot wander outside the requested labels. The result still comes from a model, so a valid shape does not make the decision correct.
The API mirrors TypeSafe Jev closely enough for its SDKs to point at an OpenJev server. Model aliases cover the normal SDK defaults, while pinned Jev versions do not. OpenJev also adds images, repeated samples, extra denoising steps, and an optional thought before the read. Those extensions make the server more flexible, but the basic appeal is narrower: 3 question types arrive with probabilities and a confidence derived from the distribution.
Five model routes have different limits
DiffusionGemma is the headline backend, a 26B model with 4B active parameters that accepts text and images. The same server can route to Laya, Verdict, CLM, or JevK5. Those alternatives range from a 151M encoder to an 8B base model with contrastive heads. They also change what an input means: some truncate long text, some reject it, and only DiffusionGemma supports images, extra steps, samples, and think.
This is useful choice, not interchangeability. Verdict ignores custom criteria on a yes-or-no question. Laya shares 256 option tokens and the README suggests keeping choices to about 20. The CLM section warns that score questions can ignore the state. A team should select one model for one decision class and test it, rather than treating the 74-package Python install as proof that every listed backend behaves alike.
What happened when we ran it
Our sandbox installed commit 4e7fd24 in 29 seconds. It pulled 74 packages and occupied 219 MB on disk, then completed the build in 6 seconds. Pytest finished in 30 seconds with 98 passed, 0 failed, and 40 skipped. Pip-audit reported 0 known vulnerabilities. The checkout itself contained 29 files and about 4,554 lines of source.
Those numbers cover repository health, not inference quality or GPU service startup. We used a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. No model weights were loaded. The scan found a tests directory and a Compose file, but 0 CI workflow files and no root Dockerfile. Backend-specific Dockerfiles live under docker/, as the README explains.
Self-hosting starts at 18 GB of weights
The default NVIDIA path needs at least 24 GB of GPU memory and downloads about 18 GB of weights on first start. The shared CUDA, Python, and PyTorch image layers take about 8.7 GB according to the README. Running DiffusionGemma beside both small encoders brings the documented storage total to about 21 GB. This is an infrastructure component, even though its Python package installed in 29 seconds on our box.
Apple silicon is easier to try because it needs no Docker or vLLM, but the 4-bit model still takes about 16 GB just to load. The MLX backend processes reads one at a time and is described as local-use software, not a serving tier. A free hosted Codiv endpoint removes the hardware job and supplies 100M input tokens, according to the project, but it moves requests outside your own system and requires a service credential.
The vLLM image carries two local patches
OpenJev pins an upstream vLLM commit and changes two behaviors. One raises the exact-label limit to support choices with as many as 255 options. The other gives image tokens bidirectional attention for DiffusionGemma. The README says an image build fails if either change stops applying. That is candid documentation, and it also tells an operator exactly where an upstream update may break.
Capacity needs equal care. A wait longer than 120 seconds returns a 503, and the project's own long-context table shows that concurrency can add queueing once prompt prefill saturates the GPU. Authentication is optional until OPENJEV_API_KEY or an origin secret is configured. Anyone exposing port 8080 beyond localhost should set that before the first real request.
Recent fixes are active, while releases are untagged
The repository was pushed on October 6, 2026, and GitHub showed 628 stars plus 2 open issues and pull requests on October 7. One open item is a benchmark-results discussion and the other is a proposed model addition. Recent merged work fixed one-option confidence, bounded MLX memory, corrected score documentation, and hardened request intake. The project has no GitHub release tagged as latest, so deployers must pin a commit or the documented container tag themselves.
OpenJev earns a serious trial when your application already asks bounded questions and can act on a probability. Its typed output is a concrete improvement over repairing malformed model prose. The cost is equally concrete: model-specific semantics, a large hardware footprint for DiffusionGemma, and a patched serving stack. The 98 passing tests tell us the Python layer is in good order; only a dataset drawn from your own decisions can tell you whether its answers belong in production.

