Replay compares the log with itself, not with live tools
ToolReplay answers a narrow question: does this recorded agent session contradict or repeat itself? It canonicalizes each tool name and argument object, remembers the first response, and flags a later identical call if the recorded answer changes. The earliest conflicting call becomes the divergence index. No command, API, browser, or file tool runs during that check. The transcript is both evidence and oracle.
That choice keeps audits safe and deterministic. It also limits the claim. A six-call sample can reveal that the same search arguments produced 3 hits and later 7, but ToolReplay cannot tell which answer was correct or whether the outside world legitimately changed. Open pull request 1 addresses one concrete version of this problem: identical shell arguments may receive different piped input that the current four-field record never captures.
Three findings cover inconsistency, repetition, and tool-name scope
The replay command reports non-determinism and redundant calls. Redundancy is deliberately conservative: identical calls count only when nothing between them appears able to change state. Five built-in tool names, including write_file and run_command, are treated as mutators. Any different intervening call also prevents the warning. This catches an adjacent duplicate read without labeling every later recheck wasteful.
The separate scope command compares each recorded tool name with an exact, case-sensitive allowed_tools list. In the shipped example, a docs-reader may call read_file, list_dir, and search; its write_file call is reported as permission overreach. The rule is easy to understand and easy to automate. It cannot tell whether an allowed read_file reached outside the permitted directory, because it never evaluates arguments against policy.
What happened when we ran it
Our sandbox installed commit a9f1374 in 33 seconds on Debian with Python 3.12, 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges. The environment pulled 36 packages and occupied 37 MB. The build passed in 7 seconds. Pip-audit reported 0 known vulnerabilities in that installed environment.
Pytest completed in another 7 seconds with 34 passed and 0 failed. The supplied tests cover parsing, canonical encoding, chaining and tamper detection, replay findings, scope errors, samples, all five CLI commands, and exit codes. The repository also has one GitHub Actions workflow that repeats its unittest suite on Python 3.11 and 3.12 for pushes and pull requests.
The checkout itself was 0.1 MB, containing 35 files and about 1,037 lines of source. That size suits a tool whose rules should remain inspectable. Installation still created a 37 MB environment with 36 packages in our harness, despite the package declaring no third-party runtime dependency. Those are environment measurements, not a claim that ToolReplay imports 36 libraries when it runs.
A hash chain detects edits but cannot prove authorship
seal turns each record into a SHA-256 chain link. Every digest includes the previous digest plus the current index, tool, arguments, and response. Change the response at index 3 and verify reports index 3 as the first broken link. Reordering or editing an earlier record changes the later chain as well. Deterministic JSON encoding keeps the same input stable across runs.
This is useful for noticing that a sealed artifact changed in storage or during review. It is not a digital signature. Anyone able to rewrite the transcript can recompute every digest and create a new internally consistent file. ToolReplay says this plainly. If provenance matters, sign the final digest with a separately managed identity key and retain that signature somewhere the transcript editor cannot replace.
Strict JSONL makes automation predictable and integration manual
Each nonblank line must contain exactly four fields: a zero-based consecutive index, a string tool name, an object of arguments, and an object response. Missing fields, extra fields, wrong types, gaps, and malformed JSON stop the audit with exit code 2. Findings use exit code 1, while clean results use 0. CI can distinguish a session problem from input the auditor could not read.
Strictness prevents quiet repair, but most agent platforms will need an adapter. Text responses must become objects. Metadata such as timestamps, duration, model identity, stdin hashes, or trace IDs cannot sit beside the four accepted fields at commit a9f1374. Keep the original trace as the source record and generate ToolReplay's smaller view rather than throwing away data to make the parser happy.
Version 0.6.0 is clear about the current ceiling
The current GitHub release is v0.6.0, published September 14, 2026. GitHub records 182 stars, 21 forks, and one open pull request; its combined issues-and-PR count is 1. The repository was last pushed on September 14. That is enough activity to inspect, but not enough history to infer how it behaves across many recorder formats or long production traces.
The README names the missing pieces instead of hiding them: configurable mutators, argument-aware scope, a signature option, and machine-readable findings remain directions rather than promises. Our 34/34 result supports the implemented rules at commit a9f1374. Adopt ToolReplay when those rules match the audit you need. If your decision depends on path policy, live behavior, semantic equivalence, or trusted provenance, pair it with a recorder and policy system built for those jobs.

