LocalJev trades direct logits for portable JSON probabilities
LocalJev accepts the POST /v1/systemone shape used by Jev, translates each typed question into a classification prompt, and asks an OpenAI-compatible model for probabilities in JSON. Version 0.2 supports choices, ordered scores, and yes-or-no questions, then returns expected values and an entropy-based confidence field. That makes it a compact adapter for code already written against the TypeSafe SDK.
The adapter changes what the probability means. OpenJev uses special DiffusionGemma operations to inspect selected-token logits during a structured read. Ordinary oMLX chat endpoints do not expose those operations, so LocalJev asks the model to write its own probability values. The README is admirably direct: the result is wire-compatible, but not mathematically equivalent. If a decision affects money, access, or safety, calibration on your own labeled workload is mandatory.
The 30-file wrapper does not include the model server
commit 3f23e36 contained 30 files, about 2,136 lines of source, and a 1.4 MB checkout. The service itself needs Bun 1.2 or newer, but useful output also requires a running oMLX server and an installed checkpoint. Its defaults point to 127.0.0.1:8000 and diffusiongemma-26B-A4B-it-4bit. LocalJev does not download, load, or operate that model for you.
Configuration is plain environment variables. The API listens on port 8080, uses a 180-second upstream timeout, and allows 2 calls in flight by default. You can require a bearer token from clients, while a separate key goes to the inference server. The TypeSafe SDK still expects a key even when LocalJev authentication is disabled, so the README uses a dummy local value. These details are documented well enough to avoid a source-code hunt.
What happened when we ran it
Our fresh Debian sandbox installed commit 3f23e36 in 12 seconds. Bun added 9 packages, and the installed project occupied 31 MB. The repository has no build target, so there was no build command to run and that step was skipped. A missing build script is reasonable for a Bun service executed from TypeScript, though teams that publish compiled artifacts will need to supply that path.
bun test finished in 5 seconds with 19 passed and 0 failed. The container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges. Those tests cover repository behavior without proving a live inference call, because the runtime path needs the external server and model. Our scan also found 0 CI workflow files and no Dockerfile, so the green local suite is not backed by a visible repository gate or official container recipe.
Valid JSON can still contain a bad probability
LocalJev validates the model's entire reply, retries malformed output up to 2 times, normalizes vectors, and splits large jobs at 16 questions or 128 outcomes by default. Those checks can catch broken JSON, missing values, or distributions that need normalization. They cannot tell whether a cleanly formatted 0.91 deserves that confidence. Syntactic correctness and calibration are different tests.
The repository includes a more serious evaluation harness than its size suggests. Its published screening run used 1,200 requests across five installed models, three public datasets, and two input lengths. The accompanying guide pins dataset revisions, records failures, and warns that 40 examples per task are too few for fine ranking or calibration claims. That work is useful evidence about the method, but it is the author's M5 Max run, not a result from our Debian sandbox or proof about your data.
The waiting queue rejects work earlier than configured
The default settings advertise 2 active upstream calls and 64 waiting decisions before HTTP 529. Open issue 2 demonstrates that the engine compares active plus waiting work against the queue limit. With both limits set to 1, the second request is rejected instead of occupying the promised waiting slot. Pull request 3 contains a fix, but it was still open on October 6, 2026. Operators should treat the current queue capacity as smaller than the README says.
GitHub showed 817 stars and 3 open issues and pull requests on October 6. The split was one issue and two pull requests, not three confirmed bugs. The last push to the default branch was September 18, the day commit 3f23e36 added the evaluation framework. There are only 3 commits and no tagged release. That is enough activity to call the project new, but not enough history to infer how quickly fixes ship.
OpenJev keeps the probability path closer to the model
OpenJev is the closer substitute when you need the original structured-read approach and have a supported NVIDIA machine for its patched vLLM backend. Outlines is a better fit when the job is constrained output rather than Jev compatibility. BAML goes wider, defining typed model functions and clients for application code. Neither promises that model-written probabilities are calibrated merely because they match a schema.
LocalJev earns a trial when you already run oMLX and need the TypeSafe interface on local hardware. The 31 MB installed footprint and 19 passing tests make the adapter easy to inspect. The harder dependency sits outside the repository: model hosting, checkpoint choice, and workload-specific evaluation. Until the queue fix lands and labeled results support your use case, treat it as a careful experiment rather than decision infrastructure.

