One endpoint replaces Jev for three typed question forms
jeff exposes /v1/systemone and accepts the official TypeSafe SDK's choice, score, and noul question types. A choice selects a label, a score maps an ordered set, and a noul returns a probability for yes. Existing clients can switch by setting TYPESAFE_BASE_URL and an API key, rather than rewriting their request objects around another classifier API.
The model behind that contract is GLiFormer Large with 400 million parameters. Device selection prefers CUDA, then Apple MPS, then CPU. This is a focused substitute for bounded decisions, not a general text-generation server. The README says jeff is less accurate than hosted Jev on reasoning-heavy work, which is the right warning to keep beside the phrase "drop-in replacement." Wire compatibility does not make model behavior identical.
The 5,774 MB environment is larger than the 10.5 MB checkout
Our checkout contained 68 files, about 4,564 lines of source, and occupied 10.5 MB. Installing 124 packages expanded the environment to 5,774 MB before a production model request. You must also download a GLiFormer checkpoint. That footprint makes jeff a deliberate service deployment, even though the application code and HTTP surface are small.
Python 3.12 and uv are the documented local route. The server can run with PyTorch on CUDA, MPS, or CPU, while ONNX is an extra path that moves the encoder to ONNX Runtime and leaves the recurrent and classification layers in PyTorch. The README calls CPU a fallback and recommends MPS on a Mac. GPU hosting examples target Modal, with environment variables controlling concurrency, batching, warm containers, and limits.
What happened when we ran it
Our sandbox installed commit 34b32f9 in 51 seconds, pulling 124 packages and using 5,774 MB on disk. The build passed in 1 second. Pytest failed after 9 seconds: 22 tests passed, 2 failed, and 6 were skipped in the recorded 24-test run. Pip-audit reported 0 known vulnerabilities.
Both failures came from tests/test_sdk_live.py. test_sdk_roundtrip and test_sdk_errors each stopped with ModuleNotFoundError: No module named 'typesafe_sdk'. The log also included one Starlette deprecation warning about using httpx with starlette.testclient. It did not show a model error, server assertion, or network failure, and we will not guess why the SDK module was absent.
The README says uv sync --extra dev installs typesafe-sdk, and it describes these as live-server tests using a fake backend. Our fresh Debian environment did not complete them. The container had 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. There was a tests directory, but no CI workflow and no Dockerfile to demonstrate another reproducible path.
Question sharing can change the answer
Choice and score questions share an encoder pass by default, while nouls receive separate passes. The README warns that choice and score questions can affect one another. Setting JEFF_ISOLATE=all gives each question its own pass at added cost. Open issue 4 supplies a sharper example: reversing question order reportedly changed a choice from Billing to Technical and exposed a mismatch between the displayed probabilities and the score calculation.
That issue matters for applications that expect the same question to mean the same thing regardless of its neighbors. Before migration, replay a representative production set in the same batches and order you plan to use. Compare labels, probabilities, and threshold decisions against hosted Jev. Isolation is available, but it changes the compute tradeoff that makes a local classifier attractive.
Authentication is optional until you set a key
JEFF_API_KEYS defaults to empty, which leaves authentication off. Set one or more bearer keys before exposing the service beyond a trusted loopback or private network. The API also has per-key rate limiting, a bounded request queue, limits for question count, label count, and input length, plus health and statistics endpoints. A full queue returns HTTP 529 rather than waiting without a bound.
The response headers include request, server-time, and batch-time identifiers. Those controls are sensible for a small inference service, but deployment packaging remains unfinished. Two open pull requests cover a label-collision and serialization-limit fix, and Docker support with Compose and an environment example. Neither was merged in the state we reviewed.
The project's own benchmark favors cost over intelligence
The README reports results from 1,600 labeled items across eight datasets and JevBench v1.2.2. Its table puts jeff behind hosted Jev on AG News accuracy and the JevBench intelligence score, while estimating a lower cost per million single-question requests on an L4 Modal deployment. The text says jeff's JevBench rank comes from cost and places it 14th of 18 on intelligence.
Those are project-supplied results, not measurements from our sandbox. We did not download model weights or run the benchmark suite, so they should guide a reproduction plan rather than settle a purchase decision. Workload mix and utilization affect the cost comparison, and question isolation can alter both behavior and throughput.
September activity shows an early project, not a finished appliance
GitHub showed 295 stars and 3 open issues and pull requests on October 7, 2026. The fetched queue held one issue and two pull requests. The last push was September 20, one day after the repository was created, and there is no tagged release. Recent attention is visible, but the history is short and none of the open fixes has a release vehicle yet.
Self-hosting a compatible endpoint is a useful idea, especially when state must remain inside your environment. jeff already documents the behavioral differences instead of hiding them. Still, our run found a 5,774 MB environment and a red suite at the SDK boundary. Treat this as a model you evaluate beside Jev, not a service you swap into production on compatibility claims alone.

