mrkeyoor.com_
Mon 05 Oct 06:28 UTC
LLM Toolsevaluationupdated 05 Oct 2026

fast-jev-compaction review

fast-jev-compaction is a TypeScript library and early-access Claude Code plugin that prunes old tool calls and results instead of replacing a session with a prose summary. It asks the hosted Jev decision service what to keep, preserves user and assistant text, and falls back to Claude Code's built-in summary when the hook fails or saves too little.

Verdict

Our fast-jev-compaction run installed 53 packages in 8 seconds and passed all 29 tests, with 0 audit findings, but those tests never contacted Jev. Try the library on disposable transcripts if verbatim surviving text matters more than a summary, then inspect what each threshold removes. Do not enable the Claude Code hook across sensitive or consequential work until the open calibration and missing-evidence reports are resolved or your own replay tests show acceptable behavior.

We ran it

Lab card: what happened when we ran fast-jev-compactionScreenshot of fast-jev-compaction (github.com/tamaratran/fast-jev-compaction)
Install✓ · 8s53 packages · 64 MB
Build✓ · 7s
Tests✓ · 8s29 passed · 0 failed of 29 (vitest)
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo26 files~13,808 lines of source · 0.6 MB · 0 CI workflows · tests dir

Answers from our run

Does fast-jev-compaction build from source?

Dependencies installed in 8 seconds (53 packages), and the build succeeded in 7 seconds. We cloned commit e3f262a into a clean Debian container with 3 CPUs and no project-specific setup.

Do fast-jev-compaction's tests pass?

Yes: 29 of 29 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Does fast-jev-compaction have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use fast-jev-compaction?

Sessions containing secrets or proprietary prompts that cannot leave the machine: the README says the state includes conversation text and tool inputs and is sent to Jev, once per request batch.

What are the alternatives to fast-jev-compaction?

Claude Code built-in compaction, claude-mem, Mem0. Our fast-jev-compaction run installed 53 packages in 8 seconds and passed all 29 tests, with 0 audit findings, but those tests never contacted Jev.

Setup3/5Fast local install, but plugin flags, API credentials, and tuning remain
Docs4/5Clear algorithm and options, with privacy and failure risks too quiet
Community4/57,388 stars and heavy issue activity within weeks of creation
Maturity2/5Early hook, no release, no CI, and consequential open bug reports

Who it’s for

Claude Code users whose long sessions are dominated by old tool output rather than irreplaceable conversation text.
Agent developers who want a composable TypeScript library for pairing calls and results, fitting state, batching decisions, and applying deletions.
Teams able to inspect every compaction decision and keep a built-in summary fallback.
Experimenters comfortable with an early-access Claude Code hook and a hosted decision API.

Who it’s NOT for

Sessions containing secrets or proprietary prompts that cannot leave the machine: the README says the state includes conversation text and tool inputs and is sent to Jev, once per request batch.
Anyone who assumes a 0.5 keep threshold is universally safe: issue 56 reports that this default dropped or truncated results it ranked as important.
Agents where removing the evidence but retaining the assistant's narration could trigger false claims: issue 65 documents 9 consecutive tool-free completion reports after one compaction.
Teams requiring stable Claude Code APIs: function hooks are documented as early access from version 2.1.274.
Buyers looking for a tagged, quiet release line: there is no GitHub release, while 99 issues and pull requests were open.

Setup reality

Our fresh Debian sandbox installed commit e3f262a in 8 seconds, adding 53 packages and using 64 MB. The build passed in 7 seconds, and Vitest finished in 8 seconds with 29 passed and 0 failed. Npm audit found 0 known vulnerabilities.

Library use needs a TypeSafe API key unless you implement the JevAsker interface. Plugin use also requires Claude Code 2.1.274 or newer, the early-access function-hook flag, marketplace installation, a restart or plugin reload, and threshold choices that deserve testing on your transcripts.

The repository has a tests directory, but our scan found 0 CI workflow files and no Dockerfile. Unit tests use a fake Jev and do not contact TypeSafe, so the green suite checks local logic rather than current service behavior or compaction quality on a real session.

The plugin deletes tool evidence instead of rewriting the conversation

fast-jev-compaction takes a different bet from normal context summarization. User and assistant text stays in order and unchanged. Old tool calls and results are the candidates for removal. Each call is paired with its result, recent messages are pinned, and Jev returns separate probabilities for keeping the call and keeping its result. The plugin then keeps both, keeps the call with a shortened result, or removes the pair.

This protects exact instructions and assistant wording from a lossy summary. It also creates a sharp failure mode: an assistant sentence can survive after the tool evidence behind it disappears. The README says unmatched results are prevented and failures throw, but structural validity does not prove that the remaining history tells the truth. Issue 65 describes a session where the model produced 9 tool-free reports of work that had not happened after compaction removed calls.

A 0.5 threshold can erase results the scorer ranks highest

The default keepThreshold is 0.5 for both the call and result questions. Issue 56 reports that the two answers came back on different scales: needed result scores topped out below the default while their call scores cleared it. In the reported 183-message transcript, the default dropped all 141 non-pinned candidates. This is one contributor's measurement, not our lab result, but the supplied reproduction and tables make it a serious adoption blocker.

Lowering the number is not a general fix. That same report found a large swing in retained size across small threshold changes, and issue 123 records different behavior at 0.3 and 0.15. The right cutoff can change with the transcript and model. A setting that produces a pleasing reduction percentage may still remove the exact error, path, or completed side effect the agent needs next.

What happened when we ran it

Our sandbox installed commit e3f262a in 8 seconds, pulling 53 packages and occupying 64 MB on disk. The build succeeded in 7 seconds. Vitest then completed in 8 seconds with 29 passed and 0 failed. The container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges. Npm audit reported 0 known vulnerabilities.

Those results establish that the small TypeScript codebase installs, compiles, and passes its supplied local checks. The checkout held 26 files and about 13,808 lines of source in 0.6 MB. Our scan found a tests directory, 0 CI workflow files, and no Dockerfile. The lack of a workflow means the repository does not show GitHub enforcing the same 29-test gate on changes.

The tests deliberately use a fake Jev and never contact TypeSafe. We therefore did not measure live decision quality, request cost, latency, transcript reduction, or whether the current endpoint accepts the generated payload. A passing 8-second suite cannot answer whether deletion choices preserve enough evidence for an agent to finish its task.

Every batch sends substantial session state to Jev

The state sent to Jev includes conversation text and tool inputs, while result bodies become short notes such as a character count. When one request cannot hold all questions, the plugin sends the same fitted state again with each batch. That design gives every decision shared context, but it also makes data handling a first-order review item. An API key is only one part of that decision.

Read your sessions as potential outbound data. User prompts can hold customer information. Assistant text can repeat code or credentials, and tool inputs can contain paths, commands, or query values. The README warns against committing the key but gives less prominence to the session content sent for scoring. Issue 88 asks for clearer disclosure. Teams with sensitive repositories should not treat an omitted tool result as proof that no sensitive content leaves the machine.

Early-access hooks and open reports raise the operating cost

The plugin requires Claude Code 2.1.274 or newer and an environment flag that enables function hooks. Its fallback to built-in compaction is sensible, but it means two histories can behave differently depending on service errors or achieved reduction. Issue 107 reports unfiltered subagent and precompute events, a possible in-flight race, and a rejected request after compaction on Claude Code 2.1.282.

GitHub showed 99 open items, split into 37 issues and 62 pull requests. The default branch was last pushed on September 18, 2026, while issue and pull-request activity continued through October 5. There is no tagged release. The activity is real, but much of the proposed work has not landed on the branch our metadata check saw. Pinning a commit is safer than assuming an open fix exists in installed code.

The local library is the safer place to evaluate the idea

The package exports its pairing, fitting, batching, decision, and application stages, and callers can supply their own JevAsker. That makes the library more attractive than enabling the hook globally. You can replay copied transcripts, compare the retained history with a plain head-and-tail cut, and reject a compaction when known-needed evidence disappears.

Issue 99 reports that several selectors failed to beat a simple head-and-tail baseline meaningfully on one user's sessions. That result has limits, but it asks the right buying question: does model scoring preserve more useful evidence than a cheap deterministic cut at the same size? Until your own replays answer yes, the 29 passing tests support an experiment, not trust in automated deletion.

Alternatives

ProjectWhat it isPick it when
Claude Code built-in compactionThe default summary-based path already available inside Claude Code.pick this instead when you prefer a supported default and do not want transcript state sent to another decision service.
claude-mem gh↗A persistent memory plugin that records sessions and injects retrieved context later.pick this instead when recall across sessions matters more than pruning one live transcript in place.
Mem0 gh↗A memory layer for extracting, storing, and retrieving user or agent memories.pick this instead when your application needs an explicit memory store rather than a Claude Code compaction hook.

What people are saying

  1. [velocity-scout] tamaratran/fast-jev-compaction

Sources

  1. fast-jev-compaction README
  2. Repository metadata
  3. Issue 56: threshold calibration
  4. Issue 65: missing evidence and fabricated reports
  5. Issue 88: hook and privacy concerns
  6. Issue 99: selector replay comparison
  7. Issue 107: hook event handling

More llm tools reviews

SemIf-OpenJev · awesome-typesafe-jev · experiential · llm-master · obsidian-mind · DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks · the whole board →