SGLang's native decisions API has replaced the reason to adopt this server
The first paragraph of the README settles the buying question. openjev-sglang was an early attempt to reproduce Jev after its release, and SGLang now ships a native decisions endpoint implemented the same way. The maintainer explicitly tells readers to use that newer API. This repository still explains the method unusually well, but a clear successor exists in the engine it already depends on. New deployments should begin there.
The project itself is a Jev-compatible FastAPI service around Qwen3.6-35B-A3B on SGLang 0.5.19. It supports three answer shapes: a yes/no probability called Noul, a choice with a full distribution, and a score across ordered levels. The server validates the request, prepares one shared prompt prefix, sends a warm-up prefill, and then makes one-token calls for each question. For N questions, the backend receives N+1 calls.
One-token readout makes decisions cheap in output tokens
openjev-sglang never asks the model to write an explanation. It requests probabilities for answer labels, ignores the sampled token, and normalizes those label scores. Up to 64 answer labels are verified against the tokenizer at startup so each remains a single token. That keeps a classification request focused on prefill and first-token scoring instead of paying for an autoregressive response that must be parsed afterward.
The response still needs careful interpretation. Choice probabilities are conditioned on the options and their order. The confidence field is derived from normalized entropy, not calibrated correctness. Usage counts add full prompt tokens across warm-up and question branches, including cache hits. The README is unusually candid about those semantics, which makes the repository useful as an implementation note even after its deployment path became obsolete.
What happened when we ran it
Our fresh Debian sandbox installed commit bf6a53b in 34 seconds. The process added 71 Python packages and occupied 218 MB on disk. The build succeeded in another 6 seconds, and pip-audit reported 0 known vulnerabilities. The checkout itself was 2.1 MB, containing 83 files and about 5,014 lines of source. This is a small API layer compared with the model service it controls.
The test run failed after 13 seconds. Pytest reported 48 passed, 3 failed, 1 deselected, 1 warning, and 26 collection/setup errors across 77 tests. The three named backend failures said async functions were not natively supported. The log also listed API assertion errors for token limits, overload handling, authentication, body limits, model metadata, the root UI, and cache-count reporting. We cannot tell from the supplied tail why all 26 setup errors occurred.
That result matters because the repository has a tests directory but no CI workflow files. A passing local build proves packaging works. It does not show that the service behavior is ready for production, and there is no visible automated check in the repository to settle the difference. GitHub showed only 5 combined open issues and pull requests, which is manageable, but also reflects a project created on September 17, 2026.
The documented deployment assumes one B200 per container
The primary production path uses Modal. Each container gets one B200, imports a large SGLang image, downloads model weights on its first GPU start, and captures or compiles kernels. Persistent Modal volumes retain weights and tuning caches, but CUDA graph capture still happens at startup. A server scaled to zero returns 503 while it wakes, and the included smoke command retries those responses.
Those are sensible mechanics for an experiment, not neutral infrastructure. The sample endpoint sets unauthenticated=True, has no explicit container cap, and spans several US regions. The default scale-to-zero delay is 5 idle minutes. Before exposing it, an operator would need to add authentication, a spending boundary, capacity limits, and monitoring. An optional API key exists, but the example does not turn it on.
Text-only input and opportunistic caching narrow the fit
Chat-shaped state preserves roles and message objects, while other JSON becomes compact text. The documented route does not accept image, audio, or video content. The service exposes health, liveness, limits, models, OpenAPI, Swagger, and a Scalar reference page. Invalid schemas return 422, oversized bodies return 413, overload returns 529, and backend timeouts return 504. Those are useful interface details for anyone maintaining an existing client.
Caching has limits too. The common prefix is warmed in SGLang's radix cache, but reuse is opportunistic rather than pinned to a request. Cache pressure, page boundaries, recurrent model state, and concurrency can reduce hits. The Rust frontend used here omits cache counts, so the API leaves the header absent rather than reporting zero. That honesty is good, though it also makes per-request cache diagnosis less direct.
Active October code does not reverse the supersession notice
The repository was pushed on October 1, 2026, and GitHub showed 336 stars. Open work included Python 3.10 support, temperature fitting, multimodal input, another backend, and a JevBench result report. No release was returned by GitHub's latest-release endpoint, and GitHub did not identify a license. The project is active in the ordinary sense. It is still the wrong starting point because its own documentation names the replacement.
SGLang is the direct choice for a new decisions service. vLLM fits a broader generation endpoint, while llama.cpp fits local and varied hardware. openjev-sglang remains worth reading for its N+1 request pattern, answer-label handling, error contracts, and cache accounting. Our 29 failed or erroring test outcomes make production adoption harder to justify, and the maintainer's notice removes the remaining reason to try.

