mrkeyoor.com_
Wed 07 Oct 06:42 UTC
AI Toolsevaluationupdated 07 Oct 2026

openjev review

OpenJev is a self-hosted decision server that answers yes-or-no, multiple-choice, and scored questions with probabilities instead of free-form text. It speaks TypeSafe Jev's wire protocol and can use DiffusionGemma or four smaller open models, giving developers a typed alternative to parsing an LLM response.

Verdict

Our OpenJev run passed 98 tests with 0 failures in 30 seconds, but serving its main model still calls for an 18 GB download and at least 24 GB of NVIDIA memory. Use it when typed probabilities solve a real parsing problem and you can evaluate the chosen model against your own decisions. Choose the hosted route or a smaller encoder if operating patched vLLM images is more work than the endpoint is worth.

We ran it

Lab card: what happened when we ran openjevScreenshot of openjev (codiv.ai)
Install✓ · 29s74 packages · 219 MB
Build✓ · 6s
Tests✓ · 30s98 passed · 0 failed · 40 skipped of 98 (pytest)
Known vulns0(pip-audit)
Repo29 files~4,554 lines of source · 0.3 MB · 0 CI workflows · tests dir

Answers from our run

Does openjev build from source?

Dependencies installed in 29 seconds (74 packages), and the build succeeded in 6 seconds. We cloned commit 4e7fd24 into a clean Debian container with 3 CPUs and no project-specific setup.

Do openjev's tests pass?

Yes: 98 of 98 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does openjev have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use openjev?

Self-hosters without suitable hardware: the main backend needs an NVIDIA GPU with at least 24 GB of memory or an Apple silicon machine with about 16 GB free just to load the 4-bit weights.

What are the alternatives to openjev?

TypeSafe Jev, Laya, Verdict. Our OpenJev run passed 98 tests with 0 failures in 30 seconds, but serving its main model still calls for an 18 GB download and at least 24 GB of NVIDIA memory.

Setup2/5Package checks pass; the main model needs costly hardware
Docs5/5API, backends, memory limits, errors, and caveats are explicit
Community3/5628 stars, 2 open issues and PRs, and an October 6 push
Maturity3/598 tests passed, but 40 skipped and no CI workflows were found

Who it’s for

Teams turning text or images into routing, priority, risk, or sentiment decisions with a fixed output shape.
Jev users who want an open server behind the existing TypeSafe SDKs.
Apple silicon developers with enough unified memory for local experiments.
GPU operators prepared to test model accuracy on their own decision data.

Who it’s NOT for

Self-hosters without suitable hardware: the main backend needs an NVIDIA GPU with at least 24 GB of memory or an Apple silicon machine with about 16 GB free just to load the 4-bit weights.
Teams that need every Jev model name to work unchanged: pinned names such as jev-1.13.0 return an unknown-model error.
Buyers who need proven decision quality out of the box: the README explicitly says accuracy depends on the selected model and calls for evaluation on your own tasks.
Release processes that require upstream CI evidence: our checkout had a test directory and 98 passing tests, but the repository had 0 CI workflow files.
Long-context services expecting predictable low latency: the README's own 64K-token load test includes queued requests reaching a 120-second timeout.

Setup reality

Our fresh Debian sandbox installed commit 4e7fd24 in 29 seconds, pulling 74 packages and using 219 MB. The build passed in 6 seconds. Pytest finished in 30 seconds with 98 passed, 0 failed, and 40 skipped; pip-audit found 0 known vulnerabilities.

That successful run covered the Python project, not a working model service. DiffusionGemma needs an 18 GB weight download plus an NVIDIA GPU with at least 24 GB, or an Apple silicon backend with about 16 GB free to load its 4-bit weights. A hosted Codiv account is the lighter route but needs an API key.

The repository has Compose and backend-specific Dockerfiles, though our scan found no root Dockerfile and 0 CI workflows. The vLLM image pins an upstream commit and patches two behaviors, so rebuilding it is more involved than the passing 6-second package build suggests.

Typed probabilities replace free-form answer parsing

OpenJev turns a state and a set of questions into fixed-shape answers. A question can be yes or no, a named choice, or a score with up to 10 levels. The server reads probabilities from designated answer tokens rather than asking a language model to write JSON. That removes one familiar failure mode: the answer cannot wander outside the requested labels. The result still comes from a model, so a valid shape does not make the decision correct.

The API mirrors TypeSafe Jev closely enough for its SDKs to point at an OpenJev server. Model aliases cover the normal SDK defaults, while pinned Jev versions do not. OpenJev also adds images, repeated samples, extra denoising steps, and an optional thought before the read. Those extensions make the server more flexible, but the basic appeal is narrower: 3 question types arrive with probabilities and a confidence derived from the distribution.

Five model routes have different limits

DiffusionGemma is the headline backend, a 26B model with 4B active parameters that accepts text and images. The same server can route to Laya, Verdict, CLM, or JevK5. Those alternatives range from a 151M encoder to an 8B base model with contrastive heads. They also change what an input means: some truncate long text, some reject it, and only DiffusionGemma supports images, extra steps, samples, and think.

This is useful choice, not interchangeability. Verdict ignores custom criteria on a yes-or-no question. Laya shares 256 option tokens and the README suggests keeping choices to about 20. The CLM section warns that score questions can ignore the state. A team should select one model for one decision class and test it, rather than treating the 74-package Python install as proof that every listed backend behaves alike.

What happened when we ran it

Our sandbox installed commit 4e7fd24 in 29 seconds. It pulled 74 packages and occupied 219 MB on disk, then completed the build in 6 seconds. Pytest finished in 30 seconds with 98 passed, 0 failed, and 40 skipped. Pip-audit reported 0 known vulnerabilities. The checkout itself contained 29 files and about 4,554 lines of source.

Those numbers cover repository health, not inference quality or GPU service startup. We used a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. No model weights were loaded. The scan found a tests directory and a Compose file, but 0 CI workflow files and no root Dockerfile. Backend-specific Dockerfiles live under docker/, as the README explains.

Self-hosting starts at 18 GB of weights

The default NVIDIA path needs at least 24 GB of GPU memory and downloads about 18 GB of weights on first start. The shared CUDA, Python, and PyTorch image layers take about 8.7 GB according to the README. Running DiffusionGemma beside both small encoders brings the documented storage total to about 21 GB. This is an infrastructure component, even though its Python package installed in 29 seconds on our box.

Apple silicon is easier to try because it needs no Docker or vLLM, but the 4-bit model still takes about 16 GB just to load. The MLX backend processes reads one at a time and is described as local-use software, not a serving tier. A free hosted Codiv endpoint removes the hardware job and supplies 100M input tokens, according to the project, but it moves requests outside your own system and requires a service credential.

The vLLM image carries two local patches

OpenJev pins an upstream vLLM commit and changes two behaviors. One raises the exact-label limit to support choices with as many as 255 options. The other gives image tokens bidirectional attention for DiffusionGemma. The README says an image build fails if either change stops applying. That is candid documentation, and it also tells an operator exactly where an upstream update may break.

Capacity needs equal care. A wait longer than 120 seconds returns a 503, and the project's own long-context table shows that concurrency can add queueing once prompt prefill saturates the GPU. Authentication is optional until OPENJEV_API_KEY or an origin secret is configured. Anyone exposing port 8080 beyond localhost should set that before the first real request.

Recent fixes are active, while releases are untagged

The repository was pushed on October 6, 2026, and GitHub showed 628 stars plus 2 open issues and pull requests on October 7. One open item is a benchmark-results discussion and the other is a proposed model addition. Recent merged work fixed one-option confidence, bounded MLX memory, corrected score documentation, and hardened request intake. The project has no GitHub release tagged as latest, so deployers must pin a commit or the documented container tag themselves.

OpenJev earns a serious trial when your application already asks bounded questions and can act on a probability. Its typed output is a concrete improvement over repairing malformed model prose. The cost is equally concrete: model-specific semantics, a large hardware footprint for DiffusionGemma, and a patched serving stack. The 98 passing tests tell us the Python layer is in good order; only a dataset drawn from your own decisions can tell you whether its answers belong in production.

Alternatives

ProjectWhat it isPick it when
TypeSafe JevThe hosted service whose SDK and decision API OpenJev reproduces.pick this instead when you want the original managed endpoint and do not want to operate model hardware.
Laya gh↗A smaller typed-decision encoder that OpenJev can also serve.pick this instead when a 421M text-only model and direct Python package are enough.
VerdictA 151M text classifier built for Jev-shaped probabilistic answers.pick this instead when CPU operation and a narrow text classifier matter more than image input or long context.
CLMA contrastive decision model that scores candidate actions against a state.pick this instead when you want to work directly with CLM's training and inference code rather than a Jev-compatible server.

What people are saying

  1. [velocity-scout] TheoLeeCJ/SemIf-OpenJev
  2. [hf-trending] AlexWortega/openjev (trending model on Hugging Face)
  3. [velocity-scout] Heman10x-NGU/openJev-verdict-2.0
  4. [velocity-scout] razorback16/openjev
  5. [hackernews] OpenJev
  6. [velocity-scout] ekzhang/openjev-sglang

Sources

  1. OpenJev repository
  2. OpenJev README
  3. OpenJev changelog
  4. OpenJev benchmark discussion

More ai tools reviews

embodied-jev · underclass · minecraft-agent · laya-coreml · CometixCode · openJev-verdict-2.0 · the whole board →