mrkeyoor.com_
Tue 06 Oct 06:33 UTC
Dev Toolsevaluationupdated 06 Oct 2026

abide review

Abide turns rules from AGENTS.md, CLAUDE.md, and similar instruction files into checks that run after a coding agent edits your repository. When a model judges an edit or completed turn likely to break a rule, Abide sends a repair request back to Claude Code, Codex, OpenCode, or Pi.

Verdict

Our Abide run installed 229 packages and built in 6 seconds, but the test command exited 1 after a CLI UI failure, so version 0.0.9 belongs in a guarded trial rather than a mandatory gate. Use it when semantic repository rules keep slipping past coding agents and someone will tune the rubric. Do not rely on it to block unsafe changes, because its documented failure mode is to let the edit continue.

We ran it

Lab card: what happened when we ran abideScreenshot of abide (github.com/coldteadotai/abide)
Install✓ · 17s229 packages · 318 MB
Build✓ · 6s
Tests✗ · 44s6 passed · 0 failed of 6 (vitest)
Repo192 files~15,270 lines of source · 10.7 MB · 0 CI workflows

Answers from our run

Does abide build from source?

Dependencies installed in 17 seconds (229 packages), and the build succeeded in 6 seconds. We cloned commit 0fc600f into a clean Debian container with 3 CPUs and no project-specific setup.

Do abide's tests pass?

Yes: 6 of 6 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Who should not use abide?

Security or compliance gates that must stop a bad edit: the README says every hook exits 0, and missing credentials, network failures, and timeouts let the edit through.

What are the alternatives to abide?

Claude Code, Semgrep, ast-grep. Our Abide run installed 229 packages and built in 6 seconds, but the test command exited 1 after a CLI UI failure, so version 0.

Setup3/5Two setup commands, but keys, agent hooks, and rubric tuning remain
Docs5/5Failure behavior, data flow, costs, limits, and commands are explicit
Community3/5547 stars, an October 5 push, and four active issues and PRs
Maturity2/5Version 0.0.9 is active, but our complete test command failed

Who it’s for

Developers whose repository instructions contain judgment calls that ordinary linters cannot express.
Teams willing to tune a committed rubric and review which rules fire, stay quiet, or produce noise.
Claude Code, Codex, OpenCode, or Pi users who want feedback during the agent session instead of waiting for pull-request review.

Who it’s NOT for

Security or compliance gates that must stop a bad edit: the README says every hook exits 0, and missing credentials, network failures, and timeouts let the edit through.
Repositories that cannot send changed lines to an outside decision service: direct checks go to TypeSafe, while gateway routing and enterprise retention arrangements have different terms.
Rules that need broad codebase knowledge: Abide sends the rule and diff, and open issue 24 asks for optional repository context for checks such as reusing existing helpers.
Teams requiring a clean fresh-container test gate: our command exited 1 after one CLI UI test failed, and the live Pi repair test is skipped unless you opt into a paid run.

Setup reality

Our sandbox installed commit 0fc600f in 17 seconds, adding 229 packages and using 318 MB. The build passed in 6 seconds. Tests exited 1 after 44 seconds: schema reported 6 passed, while CLI reported 268 passed, 1 failed, and 1 skipped; the failure was in test/ui.test.tsx.

Abide requires Node 22+, pnpm for repository work, and either a TypeSafe key or Vercel AI Gateway key for checks. init writes hooks or a plugin into the selected agent's user or project configuration, while the first session compiles rules into .abide/rubric.json.

Checks fail open. No key, no network, or a timeout allows the edit and logs the miss. Direct checks send changed lines to TypeSafe under your account; a local endpoint is possible through TYPESAFE_AI_BASE_URL. The repository has monorepo workspaces, no Dockerfile, and 0 CI workflow files.

Every hook exits 0, so Abide advises rather than blocks

Abide watches edits made by Claude Code, Codex, OpenCode, and Pi, then compares each diff with rules compiled from your instruction files. A score of 0.8 or higher sends the agent a named repair request. Scores from 0.5 to 0.8 become notes, while lower scores stay quiet. This is useful during a long coding session because the feedback arrives while the agent still has the task in hand.

The important boundary is in the README: all hooks exit 0. If the TypeSafe key is missing, the network is down, or a check reaches its deadline, Abide logs the miss and lets the edit continue. Pi's guide also says checks happen after mutations and are not a pre-write security boundary. Use the product as an advisory repair loop. A security or compliance control needs a separate deterministic gate that can fail the workflow.

Four agent integrations share one compiled rubric

Version 0.0.9 installs into four hosts. Claude Code and Codex receive hook entries, OpenCode loads a plugin, and Pi uses an extension. User-level installation changes files under each agent's home configuration. Project mode writes the integration into the repository so teammates can receive it with the checkout. Codex asks the user to approve four new hook entries through /hooks.

The first session compiles AGENTS.md, CLAUDE.md, and other discovered instructions into .abide/rubric.json. Each rule records its source, question, phase, and optional file globs. Edit-phase rules inspect one change quickly; turn-phase rules judge the full activity diff. There are no built-in policy rules, which is a sensible choice. It also means a repository without useful instruction files gives Abide nothing worthwhile to enforce.

What happened when we ran it

Our fresh Debian sandbox installed commit 0fc600f in 17 seconds. Pnpm added 229 packages and the installed workspace occupied 318 MB. The repository was a 10.7 MB checkout with 192 files and about 15,270 lines of source. Its monorepo build completed successfully in 6 seconds. The scan found 0 CI workflow files and no Dockerfile.

The complete test command failed with exit status 1 after 44 seconds. The schema package reported 6 passing tests. The CLI package reported 268 passed, 1 failed, and 1 skipped across 43 test files; test/ui.test.tsx had 1 failure among its 12 cases. The log tail did not include the failed assertion, so it does not support a cause. The skipped live Pi case requires configured credentials and can make a paid model call.

The replay confirmed 11 violations across 147 turns

The author's September 18 replay judged 1,256 edits and 147 turns from 93 real Claude Code sessions. An independent reviewer confirmed 10 of 39 edit flags and 11 of 15 turn flags. That is 26% precision for individual edits and 73% for completed turns in this sample. The strongest catches involved single-use abstractions, oversized files, and duplicated logic, which are exactly the rules ordinary linters struggle to express.

Those figures are evidence for trying the approach, not a universal accuracy rate. The two repositories were chosen by the author, one was private, the rubrics were specific to each codebase, and replay did not measure whether an agent repaired a catch. The benchmark also found one miss among 20 near-threshold cleared hunks. Your rule wording and diff shape will decide whether Abide helps or nags. calibrate, tune, and report exist because that maintenance is part of the product.

Diff-only checks cannot see an existing helper elsewhere

A Jev check receives the rule and changed diff, not the whole conversation or repository. Open issue 24 gives the practical example: a rule saying "reuse existing helpers" needs to know which helpers already exist. Today such rules may become deferred or weak proxies. The proposal asks for bounded, opt-in repository context, but it was still an open issue on October 6, 2026.

File capture has its own current edge case. Open pull request 34 says abide check can silently drop an untracked filename when Git quotes non-ASCII bytes, quotation marks, or backslashes. The proposed fix switches the file list to NUL separators. Until that lands, teams with such filenames should not assume a quiet check examined every new file. .abideignore adds explicit exclusions, and files larger than 64 KiB are skipped by that ignore-file path.

October 5 activity does not make version 0.0.9 settled

GitHub showed 547 stars and 4 open issues and pull requests on October 6. The split was two issues and two pull requests. The default branch was pushed on October 5 with the CLI at 0.0.9 and schema at 0.0.4, though GitHub had no tagged release. Recent commits fixed binary-file audits and large replay logs, which shows maintenance on real edge cases.

The pace is encouraging for a project created on September 18, 2026, but two and a half weeks is a thin operating history. Open issue 25 asks for a supported way to preserve hand-maintained rubrics without the session-start compiler replacing their questions. Combined with our failed suite, that is enough reason to pin the package, inspect configuration changes, and run abide check beside your existing gates before enabling hooks for a whole team.

A green build and failed suite point to a guarded trial

Abide addresses a real weakness in coding-agent instructions: many rules need judgment, and agents can forget them during the same session. Its 6-second build, detailed data-flow documentation, and turn-level replay results make version 0.0.9 worth testing. Start with a few rules whose violations are easy for a human to confirm, then watch the report for misses and noisy questions.

Keep Semgrep, ast-grep, tests, and review gates responsible for rules they can decide exactly. Abide fits the remaining category where a model's opinion is useful but allowed to be wrong. That division also makes its fail-open design tolerable. If the check disappears during a network outage, the deterministic controls still run, and the event log tells you which agent edits never received the extra review.

Alternatives

ProjectWhat it isPick it when
Claude Code gh↗Claude Code includes native hooks that can run deterministic commands around agent actions.pick this instead when one agent host and scriptable pass-or-fail checks cover your rules.
SemgrepA static analyzer for enforcing source patterns across many languages.pick this instead when a repeatable code rule can be expressed without a model judgment.
ast-grep gh↗A structural search, lint, and rewrite tool built around syntax trees.pick this instead when the rule concerns code shape and should run locally with deterministic results.

What people are saying

  1. [velocity-scout] coldteadotai/abide

Sources

  1. Abide README
  2. Abide replay benchmark
  3. Pi extension limits and live verification
  4. Repository-context issue 24
  5. Quoted-filename fix pull request 34
  6. GitHub repository metadata

More dev tools reviews

grokbot-field-notes · dsh-echocat-skill-panel · Compositor · ccodex-sleep-state · wtfjs · RustScan · the whole board →