The plugin deletes tool evidence instead of rewriting the conversation
fast-jev-compaction takes a different bet from normal context summarization. User and assistant text stays in order and unchanged. Old tool calls and results are the candidates for removal. Each call is paired with its result, recent messages are pinned, and Jev returns separate probabilities for keeping the call and keeping its result. The plugin then keeps both, keeps the call with a shortened result, or removes the pair.
This protects exact instructions and assistant wording from a lossy summary. It also creates a sharp failure mode: an assistant sentence can survive after the tool evidence behind it disappears. The README says unmatched results are prevented and failures throw, but structural validity does not prove that the remaining history tells the truth. Issue 65 describes a session where the model produced 9 tool-free reports of work that had not happened after compaction removed calls.
A 0.5 threshold can erase results the scorer ranks highest
The default keepThreshold is 0.5 for both the call and result questions. Issue 56 reports that the two answers came back on different scales: needed result scores topped out below the default while their call scores cleared it. In the reported 183-message transcript, the default dropped all 141 non-pinned candidates. This is one contributor's measurement, not our lab result, but the supplied reproduction and tables make it a serious adoption blocker.
Lowering the number is not a general fix. That same report found a large swing in retained size across small threshold changes, and issue 123 records different behavior at 0.3 and 0.15. The right cutoff can change with the transcript and model. A setting that produces a pleasing reduction percentage may still remove the exact error, path, or completed side effect the agent needs next.
What happened when we ran it
Our sandbox installed commit e3f262a in 8 seconds, pulling 53 packages and occupying 64 MB on disk. The build succeeded in 7 seconds. Vitest then completed in 8 seconds with 29 passed and 0 failed. The container had 3 CPUs, 8 GB of RAM, Node 22, no secrets, and no elevated privileges. Npm audit reported 0 known vulnerabilities.
Those results establish that the small TypeScript codebase installs, compiles, and passes its supplied local checks. The checkout held 26 files and about 13,808 lines of source in 0.6 MB. Our scan found a tests directory, 0 CI workflow files, and no Dockerfile. The lack of a workflow means the repository does not show GitHub enforcing the same 29-test gate on changes.
The tests deliberately use a fake Jev and never contact TypeSafe. We therefore did not measure live decision quality, request cost, latency, transcript reduction, or whether the current endpoint accepts the generated payload. A passing 8-second suite cannot answer whether deletion choices preserve enough evidence for an agent to finish its task.
Every batch sends substantial session state to Jev
The state sent to Jev includes conversation text and tool inputs, while result bodies become short notes such as a character count. When one request cannot hold all questions, the plugin sends the same fitted state again with each batch. That design gives every decision shared context, but it also makes data handling a first-order review item. An API key is only one part of that decision.
Read your sessions as potential outbound data. User prompts can hold customer information. Assistant text can repeat code or credentials, and tool inputs can contain paths, commands, or query values. The README warns against committing the key but gives less prominence to the session content sent for scoring. Issue 88 asks for clearer disclosure. Teams with sensitive repositories should not treat an omitted tool result as proof that no sensitive content leaves the machine.
Early-access hooks and open reports raise the operating cost
The plugin requires Claude Code 2.1.274 or newer and an environment flag that enables function hooks. Its fallback to built-in compaction is sensible, but it means two histories can behave differently depending on service errors or achieved reduction. Issue 107 reports unfiltered subagent and precompute events, a possible in-flight race, and a rejected request after compaction on Claude Code 2.1.282.
GitHub showed 99 open items, split into 37 issues and 62 pull requests. The default branch was last pushed on September 18, 2026, while issue and pull-request activity continued through October 5. There is no tagged release. The activity is real, but much of the proposed work has not landed on the branch our metadata check saw. Pinning a commit is safer than assuming an open fix exists in installed code.
The local library is the safer place to evaluate the idea
The package exports its pairing, fitting, batching, decision, and application stages, and callers can supply their own JevAsker. That makes the library more attractive than enabling the hook globally. You can replay copied transcripts, compare the retained history with a plain head-and-tail cut, and reject a compaction when known-needed evidence disappears.
Issue 99 reports that several selectors failed to beat a simple head-and-tail baseline meaningfully on one user's sessions. That result has limits, but it asks the right buying question: does model scoring preserve more useful evidence than a cheap deterministic cut at the same size? Until your own replays answer yes, the 29 passing tests support an experiment, not trust in automated deletion.

