mrkeyoor.com_
Wed 07 Oct 14:58 UTC
Self-Hostedevaluationupdated 07 Oct 2026

jeff review

jeff is a self-hosted implementation of TypeSafe's Jev System One API, backed by the 400M-parameter GLiFormer model. Existing `typesafe-sdk` clients can point at its local endpoint for typed choice, score, and yes-or-no probability questions, trading some accuracy for control over hosting.

Verdict

Our jeff run used 5,774 MB and ended with 22 passing tests, 2 failures, and 6 skips, so this is not yet a clean drop-in deployment from a fresh container. Trial it when Jev API compatibility and self-hosting matter enough to test every production question against the hosted service. Wait if you need Docker, a green suite, or stable answer independence across question order.

We ran it

Lab card: what happened when we ran jeffScreenshot of jeff (github.com/logan-markewich/jeff)
Install✓ · 51s124 packages · 5774 MB
Build✓ · 1s
Tests✗ · 9s22 passed · 2 failed · 6 skipped of 24 (pytest)
Known vulns0(pip-audit)
Repo68 files~4,564 lines of source · 10.5 MB · 0 CI workflows · tests dir

Answers from our run

Does jeff build from source?

Dependencies installed in 51 seconds (124 packages), and the build succeeded in 1 seconds. We cloned commit 34b32f9 into a clean Debian container with 3 CPUs and no project-specific setup.

Do jeff's tests pass?

Not all of them: 22 of 24 passed and 2 failed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does jeff have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use jeff?

Release gates that require a clean fresh-environment suite: our run had 2 failures because typesafe_sdk could not be imported.

What are the alternatives to jeff?

TypeSafe Jev, SetFit, BentoML. Our jeff run used 5,774 MB and ended with 22 passing tests, 2 failures, and 6 skips, so this is not yet a clean drop-in deployment from a fresh container.

Setup2/55,774 MB installed and two SDK tests failed
Docs4/5Clear API, deployment, limits, and benchmark caveats
Community2/5295 stars with one issue and two open pull requests
Maturity2/5No release or CI, and our complete suite failed

Who it’s for

Teams already using the TypeSafe SDK that want a local, compatible decision endpoint.
Python operators with CUDA, Apple MPS, or Modal GPU capacity for a 400M-parameter classifier.
Workloads built around bounded labels, ordered scores, and binary judgments.
Engineers willing to benchmark their own questions and tune batching, isolation, and calibration.

Who it’s NOT for

Release gates that require a clean fresh-environment suite: our run had 2 failures because typesafe_sdk could not be imported.
Small deployments with tight disk budgets: 124 packages occupied 5,774 MB before model operations.
Reasoning-heavy classification where accuracy matters more than hosting control: the README says jeff trails Jev on those tasks.
Operators who need an included container setup: the repository has no Dockerfile, and Docker support is still an open pull request.
Callers assuming question order cannot affect answers: open issue 4 reports a choice flip after reversing question order.

Setup reality

Our sandbox installed commit 34b32f9 in 51 seconds, adding 124 packages and using 5,774 MB. The build passed in 1 second. Pytest failed in 9 seconds: 22 passed, 2 failed, and 6 skipped of 24; both failures were ModuleNotFoundError: No module named 'typesafe_sdk'. Pip-audit found 0 known vulnerabilities.

Running the API also needs Python 3.12, uv, downloaded GLiFormer weights, and an API key if you want authentication. CUDA, Apple MPS, and CPU are supported, with separate ONNX extras for the CPU path.

The repository has no CI workflow or Dockerfile. Model integration tests skip without local weights, and the server starts with authentication disabled when JEFF_API_KEYS is empty.

One endpoint replaces Jev for three typed question forms

jeff exposes /v1/systemone and accepts the official TypeSafe SDK's choice, score, and noul question types. A choice selects a label, a score maps an ordered set, and a noul returns a probability for yes. Existing clients can switch by setting TYPESAFE_BASE_URL and an API key, rather than rewriting their request objects around another classifier API.

The model behind that contract is GLiFormer Large with 400 million parameters. Device selection prefers CUDA, then Apple MPS, then CPU. This is a focused substitute for bounded decisions, not a general text-generation server. The README says jeff is less accurate than hosted Jev on reasoning-heavy work, which is the right warning to keep beside the phrase "drop-in replacement." Wire compatibility does not make model behavior identical.

The 5,774 MB environment is larger than the 10.5 MB checkout

Our checkout contained 68 files, about 4,564 lines of source, and occupied 10.5 MB. Installing 124 packages expanded the environment to 5,774 MB before a production model request. You must also download a GLiFormer checkpoint. That footprint makes jeff a deliberate service deployment, even though the application code and HTTP surface are small.

Python 3.12 and uv are the documented local route. The server can run with PyTorch on CUDA, MPS, or CPU, while ONNX is an extra path that moves the encoder to ONNX Runtime and leaves the recurrent and classification layers in PyTorch. The README calls CPU a fallback and recommends MPS on a Mac. GPU hosting examples target Modal, with environment variables controlling concurrency, batching, warm containers, and limits.

What happened when we ran it

Our sandbox installed commit 34b32f9 in 51 seconds, pulling 124 packages and using 5,774 MB on disk. The build passed in 1 second. Pytest failed after 9 seconds: 22 tests passed, 2 failed, and 6 were skipped in the recorded 24-test run. Pip-audit reported 0 known vulnerabilities.

Both failures came from tests/test_sdk_live.py. test_sdk_roundtrip and test_sdk_errors each stopped with ModuleNotFoundError: No module named 'typesafe_sdk'. The log also included one Starlette deprecation warning about using httpx with starlette.testclient. It did not show a model error, server assertion, or network failure, and we will not guess why the SDK module was absent.

The README says uv sync --extra dev installs typesafe-sdk, and it describes these as live-server tests using a fake backend. Our fresh Debian environment did not complete them. The container had 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges. There was a tests directory, but no CI workflow and no Dockerfile to demonstrate another reproducible path.

Question sharing can change the answer

Choice and score questions share an encoder pass by default, while nouls receive separate passes. The README warns that choice and score questions can affect one another. Setting JEFF_ISOLATE=all gives each question its own pass at added cost. Open issue 4 supplies a sharper example: reversing question order reportedly changed a choice from Billing to Technical and exposed a mismatch between the displayed probabilities and the score calculation.

That issue matters for applications that expect the same question to mean the same thing regardless of its neighbors. Before migration, replay a representative production set in the same batches and order you plan to use. Compare labels, probabilities, and threshold decisions against hosted Jev. Isolation is available, but it changes the compute tradeoff that makes a local classifier attractive.

Authentication is optional until you set a key

JEFF_API_KEYS defaults to empty, which leaves authentication off. Set one or more bearer keys before exposing the service beyond a trusted loopback or private network. The API also has per-key rate limiting, a bounded request queue, limits for question count, label count, and input length, plus health and statistics endpoints. A full queue returns HTTP 529 rather than waiting without a bound.

The response headers include request, server-time, and batch-time identifiers. Those controls are sensible for a small inference service, but deployment packaging remains unfinished. Two open pull requests cover a label-collision and serialization-limit fix, and Docker support with Compose and an environment example. Neither was merged in the state we reviewed.

The project's own benchmark favors cost over intelligence

The README reports results from 1,600 labeled items across eight datasets and JevBench v1.2.2. Its table puts jeff behind hosted Jev on AG News accuracy and the JevBench intelligence score, while estimating a lower cost per million single-question requests on an L4 Modal deployment. The text says jeff's JevBench rank comes from cost and places it 14th of 18 on intelligence.

Those are project-supplied results, not measurements from our sandbox. We did not download model weights or run the benchmark suite, so they should guide a reproduction plan rather than settle a purchase decision. Workload mix and utilization affect the cost comparison, and question isolation can alter both behavior and throughput.

September activity shows an early project, not a finished appliance

GitHub showed 295 stars and 3 open issues and pull requests on October 7, 2026. The fetched queue held one issue and two pull requests. The last push was September 20, one day after the repository was created, and there is no tagged release. Recent attention is visible, but the history is short and none of the open fixes has a release vehicle yet.

Self-hosting a compatible endpoint is a useful idea, especially when state must remain inside your environment. jeff already documents the behavioral differences instead of hiding them. Still, our run found a 5,774 MB environment and a red suite at the SDK boundary. Treat this as a model you evaluate beside Jev, not a service you swap into production on compatibility claims alone.

Alternatives

ProjectWhat it isPick it when
TypeSafe JevThe hosted API whose wire format and decision types jeff implements.pick this instead when managed operation and stronger reasoning-task accuracy matter more than self-hosting.
SetFitA framework for training efficient text classifiers from small labeled datasets.pick this instead when you can train a task-specific classifier rather than adopt the Jev request contract.
BentoMLA framework for packaging and operating custom model inference services.pick this instead when you need to design and serve several model APIs, not emulate Jev.

What people are saying

  1. [hackernews] Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
  2. [velocity-scout] firelex/jeff
  3. [velocity-scout] logan-markewich/jeff

Sources

  1. jeff README
  2. jeff repository
  3. Question-order issue
  4. Open Docker support pull request

More self-hosted reviews

sparkDash · workbuddy2api-panel · OpenGFW · splash · quivr · bindery · the whole board →