The fixture decides what the grade can mean
An iFixAi run probes whether an agent follows the rules of its actual job. The catalog covers fabricated tool use, privilege escalation, prompt injection, audit trails, silent failure, long-run drift, human escalation, and other governance behavior. That is more useful than asking a bare model generic questions when the production system also has retrieval, tools, permissions, and policy code.
The important object is the fixture. It describes the organization, roles, expected actions, references, and ground truth that the agent should obey. The methodology warns that governance hooks are often declared in a fixture instead of measured. It also says scores are comparable only on the same fixture and release. A precise A-F grade can still rest on a weak description of the job.
Independent judging needs another provider
The system under test and the judge have separate roles. With credentials for 2 providers, iFixAi can choose a judge from a vendor other than the system being graded. With one credential, the user must explicitly select self-judging, and the scorecard carries a bias warning. This is a sound constraint: an agent should not be the sole authority on whether its own answer passed.
That independence costs money and configuration time. The README estimates roughly 2,000 judge calls for a full-suite run and gives paid model combinations as examples. The tested agent's usage is billed separately. Teams should start with a smaller suite, inspect the dry-run cost, and reserve a full run for a fixture that has already survived human review. More judge calls cannot repair an incorrect expected answer.
What happened when we ran it
Our sandbox cloned commit a978699 into an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Installation completed in 49 seconds, adding 129 packages and consuming 492 MB. The package build passed in 10 seconds. Pip-audit reported 0 known vulnerabilities.
Pytest failed after 7 seconds with exit code 5. Its complete result was 0 passed, 0 failed, and 0 collected; the final line said no tests ran in 0.06s. The repository had 569 files and about 101,430 lines of source, with 3 CI workflow files but no tests directory. We did not run a provider-backed agent inspection because the sandbox had no secrets.
The failure does not show a broken assertion. It shows that the standard pytest command found nothing to execute at commit a978699. The pyproject.toml defines unit, integration, and acceptance markers, which makes the empty collection more surprising. A CI badge and package build cannot substitute for runnable regression coverage in the checkout we measured.
An evaluator without discovered tests needs extra skepticism
Scoring code has many ways to be subtly wrong: a partial result can be treated as a pass, missing evidence can enter the denominator, or concurrent judge calls can exceed a budget. The open pull-request queue on September 28 included proposed fixes for those exact classes of problem, including unscored results, score inflation, consistency caps, judge-call budgets, and confidence intervals. Those proposals show useful scrutiny, but open code is not released code.
This matters because a letter grade compresses many decisions into one symbol. The scoring documentation publishes weights, mandatory minimums, confidence handling, and evidence states, which makes review possible. Before using the grade in a release gate, validate several known fixtures by hand and keep the per-inspection evidence beside the final score.
Real-agent coverage depends on adapter hooks
An OpenAI-compatible HTTP endpoint is the shortest route to testing a deployed agent. Other systems implement ChatProvider.send_message and optional hooks for tools, audit trails, authorization, retrieval, or governance. The more evidence those hooks expose, the more inspections can reach a decision. Missing evidence becomes insufficient_evidence or inconclusive, according to the methodology, rather than an automatic failure.
That behavior is honest, yet it creates a comparison trap. Two agents can expose different evidence and receive scorecards with different effective coverage. Version 4.0.0 added 10 exploratory inspections that are reported without changing the headline grade, while core categories still determine the letter. Compare the evidence set, fixture, judge, release, and coverage before comparing two letters.
Scorecards contain the conversations they assess
The security policy says scorecards store full model inputs and outputs. Resume checkpoints in bridge modes hold the same content, use owner-only file permissions, and remain after an interrupted run so work can resume. Those files may contain user text or retrieved material that should not enter source control, shared artifact storage, or a public report.
Telemetry is narrower: a pseudonymous install ID, start and completion events, version, operating system, interface, install source, and timestamp. It is off in CI and supports several opt-out methods. The policy says those events are retained indefinitely until the user requests deletion. This distinction is worth keeping clear: telemetry excludes prompts, while local scorecards intentionally preserve them.
September activity is high, but every open item was a PR
GitHub showed 16,760 stars, a last push on September 30, 2026, and release v4.0.0 on September 15. Its 22 open issues and pull requests were all pull requests when queried, so calling them 22 bugs would be wrong. The rapid patch queue and same-day repository activity show active work, while the beta classifier in pyproject.toml matches the pace of scoring changes.
iFixAi is worth studying for its fixed vocabulary of governance failures and its insistence on an independent judge. The grade should begin a review, not end one. Our 49-second install and 10-second build show that the package is approachable; the empty 7-second pytest run leaves the evaluator itself without the regression proof we expected. Keep humans on the fixture, the evidence, and every decision that turns a score into policy.

