mrkeyoor.com_
Tue 06 Oct 06:35 UTC
LLM Toolsevaluationupdated 06 Oct 2026

openjev-sglang review

openjev-sglang is a small Python server that copies the Jev decision API on top of Qwen3.6-35B-A3B and SGLang. It scores yes/no, choice, and ordered-score questions by reading first-token probabilities, but its own README now directs new users to SGLang's native decisions endpoint.

Verdict

Our openjev-sglang run installed 71 packages and built successfully, but its 77-test suite ended with 3 failures and 26 errors, and the maintainer now points new users to SGLang's native API. Do not start a new production deployment on this repository. Keep it only as a readable reference or to support an existing Jev-compatible endpoint while you plan a move to SGLang.

We ran it

Lab card: what happened when we ran openjev-sglangScreenshot of openjev-sglang (ekzhang--openjev-sglang-openjev.us-west.modal.direct)
Install✓ · 34s71 packages · 218 MB
Build✓ · 6s
Tests✗ · 13s48 passed · 3 failed · 26 errors of 77 (pytest)
Known vulns0(pip-audit)
Repo83 files~5,014 lines of source · 2.1 MB · 0 CI workflows · tests dir

Answers from our run

Does openjev-sglang build from source?

Dependencies installed in 34 seconds (71 packages), and the build succeeded in 6 seconds. We cloned commit bf6a53b into a clean Debian container with 3 CPUs and no project-specific setup.

Do openjev-sglang's tests pass?

Not all of them: 48 of 77 passed and 3 failed when we ran the project's own test command (pytest), with 26 collection errors. Some failures need services or credentials a bare container does not have.

Does openjev-sglang have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use openjev-sglang?

New production adopters: the README calls this an early experiment and tells readers to use SGLang's native decisions API instead.

What are the alternatives to openjev-sglang?

SGLang, vLLM, llama.cpp. Our openjev-sglang run installed 71 packages and built successfully, but its 77-test suite ended with 3 failures and 26 errors, and the maintainer now points new users to SGLang's native API.

Setup2/5Local install passed, but real inference needs Modal and a B200
Docs5/5Detailed request, limits, inference, caching, and deployment notes
Community2/5336 stars and 5 open issues or PRs in a very new repo
Maturity1/5Superseded experiment with 29 failing or erroring test outcomes

Who it’s for

Engineers maintaining an existing deployment of this exact Jev-compatible API.
Researchers studying prefill-only classification and one-token probability readout.
SGLang users who want a compact reference implementation for shared-prefix decision requests.
Teams that can operate a B200-backed Modal service and accept an experimental codebase.

Who it’s NOT for

New production adopters: the README calls this an early experiment and tells readers to use SGLang's native decisions API instead.
Teams that require a green supplied test suite: our run ended with 3 failures and 26 collection/setup errors among 77 tests.
Services that need image, audio, or video input: the documented chat state accepts text only.
Cost-sensitive deployments without B200 access: the provided Modal setup assigns one B200 to each container and performs model download plus kernel capture at startup.
Public API operators who expect safe defaults: the sample Modal endpoint is unauthenticated and has no explicit container cap.
Organizations that require an explicit open-source license: GitHub did not report one for the repository.

Setup reality

Our sandbox installed commit bf6a53b in 34 seconds, adding 71 packages and using 218 MB. The build passed in 6 seconds. Pytest then failed in 13 seconds: 48 passed, 3 failed, and 26 collection/setup errors were reported across 77 tests.

Local API development uses Python 3.11 through 3.13 and uv. Real inference needs a matching SGLang backend or the included Modal deployment, model weights, a B200 in the documented profile, and Modal authentication for deployment. API and backend bearer keys are optional, not enabled by default.

The first remote start downloads weights and captures kernels, while a scaled-to-zero server returns 503 during startup. Our scan found no Dockerfile and no CI workflow files, so the repository does not show an automated path that reproduces its deployment or test claims.

SGLang's native decisions API has replaced the reason to adopt this server

The first paragraph of the README settles the buying question. openjev-sglang was an early attempt to reproduce Jev after its release, and SGLang now ships a native decisions endpoint implemented the same way. The maintainer explicitly tells readers to use that newer API. This repository still explains the method unusually well, but a clear successor exists in the engine it already depends on. New deployments should begin there.

The project itself is a Jev-compatible FastAPI service around Qwen3.6-35B-A3B on SGLang 0.5.19. It supports three answer shapes: a yes/no probability called Noul, a choice with a full distribution, and a score across ordered levels. The server validates the request, prepares one shared prompt prefix, sends a warm-up prefill, and then makes one-token calls for each question. For N questions, the backend receives N+1 calls.

One-token readout makes decisions cheap in output tokens

openjev-sglang never asks the model to write an explanation. It requests probabilities for answer labels, ignores the sampled token, and normalizes those label scores. Up to 64 answer labels are verified against the tokenizer at startup so each remains a single token. That keeps a classification request focused on prefill and first-token scoring instead of paying for an autoregressive response that must be parsed afterward.

The response still needs careful interpretation. Choice probabilities are conditioned on the options and their order. The confidence field is derived from normalized entropy, not calibrated correctness. Usage counts add full prompt tokens across warm-up and question branches, including cache hits. The README is unusually candid about those semantics, which makes the repository useful as an implementation note even after its deployment path became obsolete.

What happened when we ran it

Our fresh Debian sandbox installed commit bf6a53b in 34 seconds. The process added 71 Python packages and occupied 218 MB on disk. The build succeeded in another 6 seconds, and pip-audit reported 0 known vulnerabilities. The checkout itself was 2.1 MB, containing 83 files and about 5,014 lines of source. This is a small API layer compared with the model service it controls.

The test run failed after 13 seconds. Pytest reported 48 passed, 3 failed, 1 deselected, 1 warning, and 26 collection/setup errors across 77 tests. The three named backend failures said async functions were not natively supported. The log also listed API assertion errors for token limits, overload handling, authentication, body limits, model metadata, the root UI, and cache-count reporting. We cannot tell from the supplied tail why all 26 setup errors occurred.

That result matters because the repository has a tests directory but no CI workflow files. A passing local build proves packaging works. It does not show that the service behavior is ready for production, and there is no visible automated check in the repository to settle the difference. GitHub showed only 5 combined open issues and pull requests, which is manageable, but also reflects a project created on September 17, 2026.

The documented deployment assumes one B200 per container

The primary production path uses Modal. Each container gets one B200, imports a large SGLang image, downloads model weights on its first GPU start, and captures or compiles kernels. Persistent Modal volumes retain weights and tuning caches, but CUDA graph capture still happens at startup. A server scaled to zero returns 503 while it wakes, and the included smoke command retries those responses.

Those are sensible mechanics for an experiment, not neutral infrastructure. The sample endpoint sets unauthenticated=True, has no explicit container cap, and spans several US regions. The default scale-to-zero delay is 5 idle minutes. Before exposing it, an operator would need to add authentication, a spending boundary, capacity limits, and monitoring. An optional API key exists, but the example does not turn it on.

Text-only input and opportunistic caching narrow the fit

Chat-shaped state preserves roles and message objects, while other JSON becomes compact text. The documented route does not accept image, audio, or video content. The service exposes health, liveness, limits, models, OpenAPI, Swagger, and a Scalar reference page. Invalid schemas return 422, oversized bodies return 413, overload returns 529, and backend timeouts return 504. Those are useful interface details for anyone maintaining an existing client.

Caching has limits too. The common prefix is warmed in SGLang's radix cache, but reuse is opportunistic rather than pinned to a request. Cache pressure, page boundaries, recurrent model state, and concurrency can reduce hits. The Rust frontend used here omits cache counts, so the API leaves the header absent rather than reporting zero. That honesty is good, though it also makes per-request cache diagnosis less direct.

Active October code does not reverse the supersession notice

The repository was pushed on October 1, 2026, and GitHub showed 336 stars. Open work included Python 3.10 support, temperature fitting, multimodal input, another backend, and a JevBench result report. No release was returned by GitHub's latest-release endpoint, and GitHub did not identify a license. The project is active in the ordinary sense. It is still the wrong starting point because its own documentation names the replacement.

SGLang is the direct choice for a new decisions service. vLLM fits a broader generation endpoint, while llama.cpp fits local and varied hardware. openjev-sglang remains worth reading for its N+1 request pattern, answer-label handling, error contracts, and cache accounting. Our 29 failed or erroring test outcomes make production adoption harder to justify, and the maintainer's notice removes the remaining reason to try.

Alternatives

ProjectWhat it isPick it when
SGLang gh↗The inference engine that now includes the native decisions endpoint this experiment anticipated.pick this instead for every new deployment, as openjev-sglang's maintainer explicitly recommends.
vLLM gh↗A widely used server for high-throughput language-model inference.pick this instead when general OpenAI-compatible generation matters more than Jev API compatibility.
llama.cpp gh↗A local inference runtime built around quantized models and broad hardware support.pick this instead when local deployment on varied hardware matters more than the supplied B200 path.

What people are saying

  1. [velocity-scout] ekzhang/openjev-sglang

Sources

  1. openjev-sglang README
  2. openjev-sglang repository facts
  3. SGLang decision models documentation
  4. openjev-sglang issues and pull requests

More llm tools reviews

simple-jev · kev · OptMem · SemIf-OpenJev · awesome-typesafe-jev · fast-jev-compaction · the whole board →