RepoPilot lets test evidence veto the coding agent
A bug report alone does not authorize a patch. RepoPilot first pins the relevant commits, reads trusted repository rules, asks Codex to design a reproduction, and runs that test against controlled snapshots. A repair becomes eligible only when the result fits the expected base-versus-head pattern and stays stable on repetition. The controller can then publish a separate branch and draft pull request, never merge it.
That sequence is the product. Many coding agents can edit a file and run a command. RepoPilot keeps generated tests frozen before repair, checks original and generated case identities, rejects zero-test or all-skipped runs, and records failure fingerprints. New behavior must point back to an exact requirement in the issue or pull-request description. Ambiguous results stop for review instead of becoming a model-authored excuse.
Version 1.5.0 expands that model into four optional Codex roles: Planner, Tester, Developer, and Reviewer. They share controller-owned call and token accounting, while role handoffs and goal states survive interruption. Post-merge checks can observe a merge commit and issue state, but the README draws a firm line around current capability: unattended product development, merging, and deployment are outside the project.
The controller holds credentials and publication authority
RepoPilot separates the part that reasons from the part that can act. The agent container receives bounded repository context and an OpenAI key, but no GitHub token or writable host checkout. The controller owns configuration, snapshots, task records, and GitHub publication. A separate test container runs target code with no network by default, a read-only source bind, dropped capabilities, resource limits, and a disposable workspace.
Publishing is opt-in. The controller rechecks source revisions, applies validated replacements to independent snapshots, preserves executable bits, and refuses force pushes. Existing branches or pull requests are reused only when their parent and tree match the verified result. These controls reduce accidental authority, though they do not prove that generated tests express the maintainer's actual intent. Human review still decides whether the proposed behavior belongs in the product.
The policy layer is useful even without repair. A base-branch .repopilot/policy.json can forbid text or direct calls, require calls, and define expiring path-specific exceptions. Nested AGENTS.md files feed semantic review. The documented AST checks do not resolve aliases or whole-program behavior, and prose conflicts need human interpretation, so treat policy as a focused gate rather than a complete static analyzer.
What happened when we ran it
Our Node 22 sandbox installed commit 6a5597d in 56 seconds. Npm added 21 packages, and the environment used 400 MB on disk. The TypeScript build succeeded in 14 seconds. The test step took 226 seconds and reported 180 passed with 0 failed. Npm audit found 0 known vulnerabilities across all severity levels.
The checkout contained 138 files, about 7,504 lines of source, and occupied 0.6 MB before installation. It has 2 CI workflow files and a tests directory. The scan reported no default Dockerfile, though the repository includes Dockerfile.agent and the README gives the exact build command for it. Those measurements cover the project checks, not a live Codex call, Docker repair run, or GitHub publication.
A 226-second green suite is meaningful here because verification is the central promise. The repository's own verification document says its tests use mocked agents, APIs, runners, and synthetic repositories rather than real model credits or a live repair PR. Our result therefore supports the controller logic and local fixtures. It does not establish that every target repository, Docker service, or GitHub permission setup will work.
The worker must be treated as a security boundary
The security guide recommends a dedicated development machine or VM without production credentials. Containers share the worker kernel, and configured dependency services can open an internal Docker network. The agent needs network access for model calls and receives repository text, so sensitive source handling also depends on outbound network policy. Prompt instructions and SDK sandbox settings are described as defense in depth, not a guarantee against injection.
Local reports and snapshots may contain proprietary code and logs, and the preview does not claim complete secret redaction. That warning sits awkwardly beside the README's public-repository scope, but it is the right warning for any future private use. Test code also shares a container with its reporter, which means a malicious repository could tamper with the evidence process. RepoPilot is a cautious automation controller, not hostile-code attestation.
Hard limits keep the surface finite. Public text repositories may contain at most 10,000 files or 16 MiB. Symlinks, submodules, unsupported binaries, case collisions, and path traversal fail closed. Fork pull requests, browser E2E, automatic dependency installation, webhooks, dashboards, and automatic merges are absent. One controller works serially, which favors a small maintainer queue over a high-volume hosted service.
v1.5.0 is packaged, but still a developer preview
GitHub showed 149 stars and no open issues or pull requests on October 6, 2026. The last push was September 30, nine days after v1.5.0 shipped with portable archives for Linux, Windows, and macOS plus a matching Linux agent image. The release archives are not code-signed or notarized, and Git plus Docker remain external requirements.
RepoPilot is worth a trial when your team already has reliable tests, explicit issue selection, and a spare worker VM. Its strongest idea is mundane and correct: model output is a proposal, while the controller and independent evidence decide what can be published. Aider fits hands-on pairing better, OpenHands covers broader agent work, and plain Codex is simpler when you do not need this persistent issue-to-PR machinery.

