mrkeyoor.com_
Sat 03 Oct 06:48 UTC
Dev Toolsevaluationupdated 03 Oct 2026

ToolReplay review

ToolReplay is a local command-line auditor for JSON Lines logs of AI agent tool calls. It finds repeated calls with conflicting recorded answers, obvious redundant calls, and tool names outside a declared allow-list; it can also hash-chain a transcript so later edits are detectable.

Verdict

Our ToolReplay run installed 36 packages in 33 seconds and passed all 34 tests in 7 seconds, making it the cleanest measured setup in this batch. Use it as a strict transcript linter and tamper check when your recorder can emit its exact JSONL format. Do not mistake its replay command for live tool execution or its name-based scope file for a complete permission system.

We ran it

Lab card: what happened when we ran ToolReplayScreenshot of ToolReplay (quartzjer.github.io/pennybank)
Install✓ · 33s36 packages · 37 MB
Build✓ · 7s
Tests✓ · 7s34 passed · 0 failed of 34 (pytest)
Known vulns0(pip-audit)
Repo35 files~1,037 lines of source · 0.1 MB · 1 CI workflows · tests dir

Answers from our run

Does ToolReplay build from source?

Dependencies installed in 33 seconds (36 packages), and the build succeeded in 7 seconds. We cloned commit a9f1374 into a clean Debian container with 3 CPUs and no project-specific setup.

Do ToolReplay's tests pass?

Yes: 34 of 34 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does ToolReplay have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use ToolReplay?

Teams seeking live re-execution: replay compares recorded responses inside one transcript and never calls the original tools.

What are the alternatives to ToolReplay?

AgentOps, Langfuse, Phoenix. Our ToolReplay run installed 36 packages in 33 seconds and passed all 34 tests in 7 seconds, making it the cleanest measured setup in this batch.

Setup5/533-second install, no service or credentials, and all 34 tests passed
Docs5/5Every command, rule, exit code, sample, and limitation is explained
Community2/5182 stars, 21 forks, one open PR, and a short public history
Maturity3/5v0.6.0 has CI and passing tests, but the audit model stays narrow

Who it’s for

Agent developers who already record every tool name, argument object, and response object in order.
CI maintainers who want deterministic text reports and separate exit codes for findings versus malformed input.
Security reviewers who need a small first-pass check for calls outside a tool-name allow-list.
Teams that want an offline transcript check with no network access or runtime dependencies.

Who it’s NOT for

Teams seeking live re-execution: replay compares recorded responses inside one transcript and never calls the original tools.
Permission systems that restrict paths, hosts, commands, or argument values: scope checks exact tool names only.
Anyone needing signed provenance: the SHA-256 chain detects edits after sealing, but an editor can change the file and seal it again.
Custom agent logs that cannot be converted to the strict four-field JSONL schema of index, tool, args, and response.
Audits with custom mutating tools: the fixed mutator list can treat an unknown state-changing tool as read-only and misclassify a later repeat.

Setup reality

Our sandbox installed commit a9f1374 in 33 seconds, pulling 36 packages and using 37 MB. The build succeeded in 7 seconds. Pytest finished in 7 seconds with all 34 tests passing and 0 failures. Pip-audit reported 0 known vulnerabilities. The checkout was 0.1 MB, with 35 files and about 1,037 lines of source.

ToolReplay requires Python 3.11 or newer and no credentials, database, service, or network connection. The package declares no third-party runtime dependencies; pip install . adds the toolreplay command. Its input must already exist as a strict JSONL transcript, with a separate JSON scope file for permission checks.

There is no Dockerfile because the tool is a local CLI. GitHub Actions tests Python 3.11 and 3.12. Integration effort sits outside installation: your agent recorder must capture ordered object-valued arguments and responses without extra fields, then preserve the log for audit.

Replay compares the log with itself, not with live tools

ToolReplay answers a narrow question: does this recorded agent session contradict or repeat itself? It canonicalizes each tool name and argument object, remembers the first response, and flags a later identical call if the recorded answer changes. The earliest conflicting call becomes the divergence index. No command, API, browser, or file tool runs during that check. The transcript is both evidence and oracle.

That choice keeps audits safe and deterministic. It also limits the claim. A six-call sample can reveal that the same search arguments produced 3 hits and later 7, but ToolReplay cannot tell which answer was correct or whether the outside world legitimately changed. Open pull request 1 addresses one concrete version of this problem: identical shell arguments may receive different piped input that the current four-field record never captures.

Three findings cover inconsistency, repetition, and tool-name scope

The replay command reports non-determinism and redundant calls. Redundancy is deliberately conservative: identical calls count only when nothing between them appears able to change state. Five built-in tool names, including write_file and run_command, are treated as mutators. Any different intervening call also prevents the warning. This catches an adjacent duplicate read without labeling every later recheck wasteful.

The separate scope command compares each recorded tool name with an exact, case-sensitive allowed_tools list. In the shipped example, a docs-reader may call read_file, list_dir, and search; its write_file call is reported as permission overreach. The rule is easy to understand and easy to automate. It cannot tell whether an allowed read_file reached outside the permitted directory, because it never evaluates arguments against policy.

What happened when we ran it

Our sandbox installed commit a9f1374 in 33 seconds on Debian with Python 3.12, 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges. The environment pulled 36 packages and occupied 37 MB. The build passed in 7 seconds. Pip-audit reported 0 known vulnerabilities in that installed environment.

Pytest completed in another 7 seconds with 34 passed and 0 failed. The supplied tests cover parsing, canonical encoding, chaining and tamper detection, replay findings, scope errors, samples, all five CLI commands, and exit codes. The repository also has one GitHub Actions workflow that repeats its unittest suite on Python 3.11 and 3.12 for pushes and pull requests.

The checkout itself was 0.1 MB, containing 35 files and about 1,037 lines of source. That size suits a tool whose rules should remain inspectable. Installation still created a 37 MB environment with 36 packages in our harness, despite the package declaring no third-party runtime dependency. Those are environment measurements, not a claim that ToolReplay imports 36 libraries when it runs.

A hash chain detects edits but cannot prove authorship

seal turns each record into a SHA-256 chain link. Every digest includes the previous digest plus the current index, tool, arguments, and response. Change the response at index 3 and verify reports index 3 as the first broken link. Reordering or editing an earlier record changes the later chain as well. Deterministic JSON encoding keeps the same input stable across runs.

This is useful for noticing that a sealed artifact changed in storage or during review. It is not a digital signature. Anyone able to rewrite the transcript can recompute every digest and create a new internally consistent file. ToolReplay says this plainly. If provenance matters, sign the final digest with a separately managed identity key and retain that signature somewhere the transcript editor cannot replace.

Strict JSONL makes automation predictable and integration manual

Each nonblank line must contain exactly four fields: a zero-based consecutive index, a string tool name, an object of arguments, and an object response. Missing fields, extra fields, wrong types, gaps, and malformed JSON stop the audit with exit code 2. Findings use exit code 1, while clean results use 0. CI can distinguish a session problem from input the auditor could not read.

Strictness prevents quiet repair, but most agent platforms will need an adapter. Text responses must become objects. Metadata such as timestamps, duration, model identity, stdin hashes, or trace IDs cannot sit beside the four accepted fields at commit a9f1374. Keep the original trace as the source record and generate ToolReplay's smaller view rather than throwing away data to make the parser happy.

Version 0.6.0 is clear about the current ceiling

The current GitHub release is v0.6.0, published September 14, 2026. GitHub records 182 stars, 21 forks, and one open pull request; its combined issues-and-PR count is 1. The repository was last pushed on September 14. That is enough activity to inspect, but not enough history to infer how it behaves across many recorder formats or long production traces.

The README names the missing pieces instead of hiding them: configurable mutators, argument-aware scope, a signature option, and machine-readable findings remain directions rather than promises. Our 34/34 result supports the implemented rules at commit a9f1374. Adopt ToolReplay when those rules match the audit you need. If your decision depends on path policy, live behavior, semantic equivalence, or trusted provenance, pair it with a recorder and policy system built for those jobs.

Alternatives

ProjectWhat it isPick it when
AgentOpsAn agent observability SDK and service for tracing sessions, costs, errors, and evaluations.pick this instead when automatic instrumentation and a browsable monitoring platform matter more than offline deterministic files.
Langfuse gh↗A self-hostable tracing, prompt, evaluation, and metrics platform for language-model applications.pick this instead when you need shared dashboards, datasets, scoring, and production traces across a team.
PhoenixAn open-source observability and evaluation system for AI applications and agent traces.pick this instead when trace visualization, evaluations, and OpenTelemetry integration outweigh a tiny dependency-free CLI.

What people are saying

  1. [velocity-scout] Matthew0822/ToolReplay

Sources

  1. ToolReplay repository and README
  2. Measured commit a9f1374
  3. ToolReplay v0.6.0 release
  4. ToolReplay replay implementation
  5. ToolReplay scope implementation
  6. stdin-aware replay pull request

More dev tools reviews

CUDA-for-AMD-Windows · DuoFold-Android · wutw-public · viserys-agent · birdview · YOINK · the whole board →