Open models become bounded decision endpoints without JSON generation
Simple Jev accepts shared state or chat messages plus typed questions. Choice picks from named options, Score returns an expected position on an ordered rubric, and Noul expresses support for a yes/no claim. The server builds the JSON response from selected logits, so the model never writes a completion that must be parsed. /v1/classifier is the main route, with /v1/systemone as an alias.
The implementation supports 2 to 255 Choice options, up to 50 Score levels, and a configurable ceiling of 256 questions. Those are schema ceilings, not promises that every model or GPU can fit the largest combination. Context length, repeated policy text, batch padding, cache memory, and option descriptions still consume resources. Requests beyond configured limits receive 422 instead of silent truncation.
One shared prefill reduces repeated context work inside a request
Each question shares most of its prompt with the others. The server finds the exact common token prefix, runs that prefix once, copies its KV cache, and batches the question suffixes. It then reads logits for the allowed answer labels. This avoids an autoregressive decode loop and stops each question from reprocessing the entire context independently. Cache reuse lasts for one request, not across callers.
The optimization does not make the scores truthful. Choice and Score confidence is the largest probability among allowed labels, while Noul is derived from a fixed rating-token distribution. The README says these values are not calibrated probabilities of correctness. Wording, option order, model capability, prompt policy, and precision can change results. A valid response proves the server followed its contract, not that the decision is safe.
What happened when we ran it
Our sandbox installed commit 9c11582 from the repository's RFDT/ project in 58 seconds. It added 82 Python packages and consumed 5,478 MB on disk. The build succeeded in 5 seconds. Pip-audit reported 0 known vulnerabilities. The full checkout contained 643 files, about 40,838 lines of source, and occupied 76 MB before the installed environment.
The test step failed after 15 seconds. Pytest reported 0 passed, 0 failed, and 1 collection/setup error. RFDT/tests/test_rfdt.py imported through common/prompt_builder.py into common/request_schema.py, where Python raised ModuleNotFoundError: No module named 'pydantic'. The log establishes the missing module, but not why the environment lacked it. No test body ran.
The repository has 2 CI workflow files, no Dockerfile, and a tests directory. A 5-second successful build shows the selected project can package in the lab environment. It does not repair a collection failure or validate classification quality. Anyone adopting RFDT should reproduce the exact install and test path on clean hardware before beginning a costly labeling or training run.
Model compatibility is checked, not assumed
A candidate Transformers model needs a supported implementation, working chat template, compatible cache operations, and answer labels that extend the rendered prompt by one distinct token each. The server checks label tokenization. Known Qwen and Gemma configurations receive development-selected prompt policies based on architecture details, while unknown configurations fall back to baseline with a startup warning to tune the prompt.
The included prompt search evaluates 4 policies over 477 development cases each, for 1,908 total selections. Those cases select a prompt format. They are not held-out evidence that a model generalizes. The tool pins revisions, preserves raw responses, and refuses to recommend a winner from an incomplete search. A production team still needs a disjoint evaluation drawn from its own decisions.
CPU starts small, while serious models need serious memory
The quick local command can launch Qwen3.5-0.8B on CPU with float32. Larger examples use Gemma 4 or Qwen on NVIDIA hardware, and the README warns that sparse expert activation does not remove the need to hold full model weights. The process also needs KV cache and inference buffers. CUDA or ROCm users must install the appropriate PyTorch build before the server package.
The server listens on port 8000 and offers health, model-listing, and interactive documentation routes. It can accept bounded public image URLs or base64 images for supported vision families, but rejects private-network URLs, audio, video, and tool calls. Model requests are processed serially in the current HF server. That keeps the reference implementation understandable, but limits production concurrency without another serving layer.
RFDT adds training, labeling, evaluation, and export to the same repository
Really Fancy Decision Training, or RFDT, prepares task data, obtains missing labels from a teacher, trains on allowed answer-token logits, evaluates results, and exports a student model back to the HF server. It supports LoRA and multi-GPU work. This is a different commitment from merely scoring a frozen checkpoint, and it explains part of the 5,478 MB environment our run produced.
The README says hosted fine-tuning support is part of a future Featherless rollout, while the scripts are available now. That wording matters: local RFDT exists, hosted training does not yet count as shipped. Teams should cost teacher labeling, GPU time, dataset rights, held-out evaluation, model storage, and rollback before choosing the training path.
An open conformance issue makes drop-in compatibility conditional
Open issue 8 reports that a September conformance run hit differences around the jev-latest model alias, option limits, null legends, and confidence formulas. The current README documents GET /v1/models, a 255-option maximum, and optional exact model enforcement, suggesting some surrounding behavior has changed. The issue remains open, so client compatibility should be tested against the commit and SDK version you will deploy.
GitHub showed 590 stars, 2 combined open issues and pull requests, an Apache-2.0 license, and a push on October 3, 2026. There is no published release. Activity is current, and the documentation is unusually frank about selection data, uncalibrated scores, model fit, and future hosting. Our collection failure keeps the recommendation narrow: study it, evaluate it, and fix reproducibility before depending on it.

