mrkeyoor.com_
Tue 06 Oct 06:34 UTC
LLM Toolsevaluationupdated 06 Oct 2026

kev review

Kev is a family of self-hosted decision models that read a document, answer yes-or-no, multiple-choice, or rating questions, and return probabilities instead of generated prose. Its server matches TypeSafe's System One API, so an application can move Jev-style classification and routing onto its own hardware.

Verdict

Our Kev install consumed 6,971 MB and its tests failed after 239 seconds, even though the 7-second build passed, so this is a serious ML stack rather than a small Python utility. Try Kev when fixed-choice decisions and local data control justify the model downloads, GPU planning, and threshold evaluation. Do not use it as a chat model, and do not ship automated consequential decisions from its probabilities without human review.

We ran it

Lab card: what happened when we ran kevScreenshot of kev (github.com/jaredpalmer/kev)
Install✓ · 94s105 packages · 6971 MB
Build✓ · 7s
Tests✗ · 239sran, no count parsed
Known vulns8(pip-audit)
Repo3049 files~34,069 lines of source · 852.5 MB · 1 CI workflows · tests dir

Answers from our run

Does kev build from source?

Dependencies installed in 94 seconds (105 packages), and the build succeeded in 7 seconds. We cloned commit fe64b12 into a clean Debian container with 3 CPUs and no project-specific setup.

Do kev's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does kev have known vulnerabilities in its dependencies?

pip-audit flagged 8 known advisories in the dependency tree at the time of our run.

Who should not use kev?

Chat, summarization, or open-ended answers: Kev scores supplied options and does not generate prose.

What are the alternatives to kev?

Jev, SetFit, Outlines. Our Kev install consumed 6,971 MB and its tests failed after 239 seconds, even though the 7-second build passed, so this is a serious ML stack rather than a small Python utility.

Setup2/56,971 MB installed, 8 advisories, and the test run failed
Docs5/5Detailed model cards, limits, evaluation records, and deployment paths
Community4/58,524 stars, 35 open issues or PRs, and an October 6 push
Maturity3/5Kev 1.0 is versioned, but packaging and hardware edges remain

Who it’s for

Teams routing, classifying, triaging, or checking documents against a fixed set of choices.
Jev users who want a System One-compatible endpoint under their own control.
ML engineers prepared to test probability thresholds on labelled examples from their own workload.
Apple Silicon or CUDA users choosing among the 0.8B, 4B, 9B, and 27B checkpoints.
Teams that want to fine-tune a decision model on their own categories and policies.

Who it’s NOT for

Chat, summarization, or open-ended answers: Kev scores supplied options and does not generate prose.
Python 3.14 environments: the project supports Python 3.12 and 3.13, and its dependency metadata excludes 3.14.
Long-document users who assume the server limit proves accuracy: the smaller three models accept 65,536 tokens but are validated only to 8,192.
Tool routing with Kev-0.8B: the release notes explicitly advise against that use after a below-chance result on one evaluation.
Teams requiring a clean dependency audit and green full suite before adoption: our run found 8 known vulnerabilities and the tests ended with exit 137.
Anyone planning to run pip install kev: open issue 159 says that name installs an unrelated ORM from PyPI.

Setup reality

Our sandbox installed 105 packages in 94 seconds at commit fe64b12 and occupied 6,971 MB. The build passed in 7 seconds. Tests ran for 239 seconds, then failed with exit 137; the log ended with a pytest timeout banner, one F, and a waiting tqdm monitor thread.

Local serving needs Python 3.12 or 3.13, uv, model weights, and suitable CUDA, ROCm, or Apple Silicon hardware. The first model run downloads both the adapter and base. The default server is open unless you set an API key.

The repository checkout was already 852.5 MB, and pip-audit found 8 known vulnerabilities in our installed environment. Kev-27B ships as 51 GB of full weights and needs far more hardware than the smaller adapters; start with Kev-4B only after checking its memory fit and your own labelled cases.

Four models answer typed questions without writing prose

Kev takes one state, which can be text or structured data, and a set of questions. A question can ask for yes or no, one choice from named options, or a score across ordered levels. The model returns a probability distribution and a selected result. It does not compose a paragraph. That makes the family relevant to ticket routing, policy checks, document triage, and other jobs where an application already knows the permitted answers.

The family has 0.8B, 4B, 9B, and 27B sizes under the Kev 1.0 release. The first three use adapters and pointer heads over Qwen3.5 base models. Kev-27B starts from Qwen3.8, fine-tunes the full backbone, and ships 51 GB of weights. All expose the same System One-compatible endpoint, which lets an existing TypeSafe Python client point at a local Kev server.

Kev-4B is the practical starting point, with an 8,192-token proof limit

The README recommends Kev-4B first. It targets a machine with one suitable GPU or a 32 GB Apple Silicon Mac, while the 27B model calls for an 80 GB data-center GPU. Smaller hardware can use Kev-0.8B, but the release notes warn against its tool-routing use. The server can also load a local checkpoint or a pinned Hugging Face revision instead of the moving latest name.

Serving limits need careful reading. Kev-0.8B, 4B, and 9B accept states up to 65,536 tokens, yet their validated context length is 8,192 tokens. The 4B model card says training states topped out at 7,552 tokens and reports a measurable accuracy drop at longer contract lengths. Kev-27B alone carries validation through the server limit. A successful HTTP 200 from a smaller model does not establish answer quality on a 50,000-token file.

What happened when we ran it

Our run at commit fe64b12 installed 105 packages in 94 seconds and used 6,971 MB on disk. The build completed successfully in 7 seconds. The fresh checkout was already large: 3,049 files, about 34,069 lines of source, and 852.5 MB before the environment was installed. Pip-audit reported 8 known vulnerabilities.

The test command did not complete cleanly. It ran for 239 seconds and exited with code 137. The final log showed pytest's timeout banner, one failure marker, and a tqdm monitor thread waiting inside Python's threading code. That tail does not identify the test name or prove why the process received exit 137, so we cannot narrow the failure to a particular module or missing system package.

The repository has one CI workflow and a tests directory, while our scan found no Dockerfile. CI deliberately runs a selected unit set without model weights or a live server, though it does download a tokenizer. Our result covers the supplied sandbox command and complete installed environment. It does not measure model accuracy, inference latency, GPU memory, or whether the narrower CI selection passes at this commit.

Installation occupies 6,971 MB before model weights arrive

The supported local path uses Python 3.12 or 3.13 and uv. Running uv sync --extra serve installs the API server and TypeSafe SDK, then the first inference downloads the selected adapter and base model. Python 3.14 is excluded because the required Torch wheels are unavailable according to the README. Open issue 159 adds another trap: pip install kev fetches an unrelated ORM, so install this project from its repository.

Hardware selection is automatic across CUDA, ROCm, Apple Silicon, and CPU, but issue 170 documents an AMD APU where Torch reports CUDA availability even though its ROCm wheel lacks the device's kernels. The result is a hard process crash after weight loading. An open pull request proposes an explicit device escape hatch. If your hardware is outside the documented paths, make a one-request smoke test part of provisioning before downloading the largest checkpoint.

One API key separates a local server from an open port

Kev binds to 127.0.0.1 by default. Binding to 0.0.0.0 exposes it to other machines, and the server accepts requests without authentication unless KEV_API_KEY is set. Self-hosting keeps document contents on your hardware only if the surrounding proxy, logs, model cache, and access controls do the same. The API returns request IDs and rejects over-limit states with HTTP 422 instead of silently cutting them.

Fine-tuning is a substantial part of the project. The repository includes direct training commands and agent skills that run a data, evaluation, training, calibration, and deployment flow on Modal. That convenience does not remove the need for held-out labels. A probability threshold should be chosen against your categories and error cost, because one fitted temperature cannot guarantee calibration after a domain change.

Kev 1.0 is active, but the package still says alpha

GitHub showed 8,524 stars and 35 open issues and pull requests on October 6, 2026, the same date as the latest push. Kev 1.0 shipped on October 1 with pinned checkpoints, checksums, model cards, and evaluation suites. The Python package metadata still marks version 0.1.0 as alpha. That split makes sense: the model family has a named release, while the serving and training code is moving quickly.

Kev is unusually candid about where its numbers stop applying. The model cards separate trained task families from unseen sources, publish context limits, and warn about English-only input, option-order sensitivity, date arithmetic, and consequential decisions. Pair that documentation with our failed 239-second test run and 8 audit findings. The right trial is one pinned model, one isolated endpoint, and a labelled slice of the exact decisions you intend to automate.

Alternatives

ProjectWhat it isPick it when
JevTypeSafe's hosted decision model and the API Kev aims to match.pick this instead when a managed endpoint matters more than running weights on your own hardware.
SetFitA few-shot text classification framework built on Sentence Transformers.pick this instead when you want a smaller conventional classifier for one labelled task rather than a shared decision API.
OutlinesA library for forcing language models to emit structured outputs.pick this instead when your application needs generated structured data rather than probabilities over fixed choices.
vLLM gh↗A high-throughput server for general language-model inference.pick this instead when text generation and broad model support matter more than Kev's System One-compatible decision endpoint.

What people are saying

  1. [hackernews] Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
  2. [velocity-scout] jaredpalmer/kev

Sources

  1. Kev repository and README
  2. Kev 1.0 release
  3. Kev-4B model card
  4. Issue 159: PyPI name collision
  5. Issue 170: ROCm device selection crash

More llm tools reviews

simple-jev · openjev-sglang · OptMem · SemIf-OpenJev · awesome-typesafe-jev · fast-jev-compaction · the whole board →