mrkeyoor.com_
Wed 30 Sept 20:33 UTC
AI Toolsevaluationupdated 27 Aug 2026

strix review

Strix is an open-source security testing agent that probes source code, web apps, and APIs, then records vulnerabilities with reproduction steps. It is meant to give developers a practical pentest loop without hiring a human tester for every code change, though it does not replace an independent assessment for high-risk systems.

+965stars / 7d
Verdict

Our Strix run passed 390 tests but failed 200 before pytest stopped, and pip-audit reported 3 known vulnerabilities. It is worth a controlled trial for security teams that understand pentesting and want an agent to widen routine coverage. Do not make it a release authority yet: the failed pinned suite and the open refusal-handling report both require human review of coverage and findings.

We ran it

Lab card: what happened when we ran strixScreenshot of strix (strix.ai)
Install✓ · 40s107 packages · 285 MB
Build✓ · 5s
Tests✗ · 32s562 passed · 200 failed of 762 (pytest)
Known vulns6(pip-audit)
Repo512 files~97,168 lines of source · 9.6 MB · 1 CI workflows · tests dir

Answers from our run

Does strix build from source?

Dependencies installed in 40 seconds (107 packages), and the build succeeded in 5 seconds. We cloned commit ae38fe7 into a clean Debian container with 3 CPUs and no project-specific setup.

Do strix's tests pass?

Not all of them: 562 of 762 passed and 200 failed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does strix have known vulnerabilities in its dependencies?

pip-audit flagged 6 known advisories in the dependency tree at the time of our run.

Who should not use strix?

Teams that cannot isolate an offensive tool: Strix runs shell commands, intercepts HTTP traffic, writes exploit code, and actively attacks the target you give it.

What are the alternatives to strix?

OWASP ZAP, Nuclei, PentAGI. Our Strix run passed 390 tests but failed 200 before pytest stopped, and pip-audit reported 3 known vulnerabilities.

Setup3/5Fast install and build, but Docker, model access, and test failures remain
Docs4/5Clear target, CI, viewer, provider, MCP, and authorization guidance
Community5/5August 27 push with active issue and pull-request discussion
Maturity3/5Useful workflow, with unresolved test and scan-completeness concerns

Discussed on

  1. hnShow HN: Strix - Open-source AI hackers for your apps102 points
  2. hnShow HN: Enter your domain and my open-source agent will hack it14 points
  3. hnShow HN: Strix – AI Cybersecurity Agent, securing your vibe coding5 points

Who it’s for

Application-security teams that want an agent to inspect code and exercise a live target.
Developers who need reproducible findings and remediation guidance inside CI.
Claude Code, Cursor, or Codex users who want packaged security-testing skills.
Teams able to provide Docker, an LLM account, and a tightly authorized test scope.

Who it’s NOT for

Teams that cannot isolate an offensive tool: Strix runs shell commands, intercepts HTTP traffic, writes exploit code, and actively attacks the target you give it.
Buyers who need a clean test baseline today: our pinned run ended with 200 failures, and the log says async tests lacked a suitable pytest plugin.
Operators routing models through a custom gateway without testing first: issue #1095 reports that setting LLM_API_BASE can stop the first model call on v1.5.3.
Anyone treating an AI report as proof of full coverage: issue #1155 reports a specialist refusal being handled as ordinary text, which can hide work the agent did not perform.
Projects without written authorization to test their target: the README warns that Strix actively probes systems and puts legal responsibility on the operator.

Setup reality

At commit 1c499c5, installation succeeded in 60 seconds with 99 packages and used 243 MB. The build succeeded in 11 seconds. Pytest ran for 123 seconds, reported 390 passed and 200 failed out of 590, then stopped at the failure cap. Pip-audit found 3 known vulnerabilities.

A real scan also needs Docker and either an LLM API key or a supported ChatGPT sign-in. The first run pulls a sandbox image. Optional web search, Postman collections, remote MCP servers, and authenticated targets add their own credentials; CI needs those values stored as secrets.

The test log repeatedly says async functions are not natively supported and suggests installing a suitable pytest plugin. It does not prove why that plugin support was absent, so we would reproduce the pinned environment before blaming the tests or packaging. Strix also executes offensive actions, which makes target scope, network boundaries, and disposable credentials part of setup rather than later hardening.

An agent with an offensive toolkit

Strix takes a more active approach than a code scanner. Give it a local directory, repository URL, web address, API contract, or Postman collection, and its agents can inspect code, browse pages, intercept HTTP requests, run shell commands, and write proof-of-concept exploits. Findings are stored under strix_runs, with severity, reproduction details, and remediation guidance. A local viewer shows the current run, agent graph, history, reports, and live steering.

That breadth is the appeal. Static analysis can point at a suspicious path; Strix is designed to exercise the path and see whether it works. It can combine source and deployed targets, focus on a pull-request diff, or run headlessly in CI. The supplied agent skills let Claude Code, Cursor, Codex, and other compatible tools start scans, read results, attempt fixes, and scan again. Strix can also consume tools from configured MCP servers.

The same breadth creates risk. This software has a shell, browser, interception proxy, and exploit runtime. It belongs in an isolated environment against a target covered by written permission. The README makes that warning explicit. A careless target URL or overly broad credentials can turn a useful security check into an incident.

What happened when we ran it

We cloned commit 1c499c5 into an unprivileged Debian container with 3 CPUs, 8 GB of RAM, no secrets, and a Python 3.12 uv image. The checkout contained 436 files, about 61,879 lines of source, and occupied 8.2 MB before installation.

Installation succeeded in 60 seconds. It installed 99 packages and consumed 243 MB on disk. The build then completed successfully in 11 seconds. Those two results support the README's claim that the Python side is approachable, although they do not cover the Docker image pull or a live LLM-backed scan.

The test result was poor. Pytest ran for 123 seconds, counted 390 passing tests and 200 failures out of 590, then stopped after reaching 200 failures. The tail repeatedly says async test functions are not supported natively and lists plugins such as anyio and pytest-asyncio. The final named failure, test_workspace_preview_honours_byte_budget, has the same message. That log tells us what pytest lacked during the run; it does not tell us whether the fault sits in dependency metadata, the test command, or our environment. Pip-audit also reported 3 known vulnerabilities.

The quick start leaves real work for the operator

The documented path is short: install the CLI, set a model name and API key, then point strix at a target. Docker must already be running, and the first scan downloads the sandbox image. Strix supports several model providers plus a ChatGPT subscription login. Local model gateways are documented through LLM_API_BASE, but issue #1095 reports that custom endpoints can fail before scanning because the API key reaches the model client twice. Anyone using an internal gateway should test the exact provider prefix and release before adding Strix to CI.

Authenticated testing needs more care than copying a username and password into an instruction. Use a disposable test account, restrict its privileges, and decide which hosts and actions are in scope. Postman and web-search integrations introduce more secrets. A remote MCP configuration can expose every tool from a server unless allowed_tools narrows the list. These are useful controls, but Strix cannot choose your blast radius for you.

CI behavior is sensible. Headless mode returns a nonzero exit code when it finds vulnerabilities, and quick pull-request scans can limit work to the diff when full Git history is available. A security team still needs a policy for model errors, contested findings, and accepted risks. Otherwise a nondeterministic agent becomes an unpredictable merge gate.

Proof of exploitation has a trust problem

A working reproduction is more useful than a long list of guesses. Strix's report model, proxy history, saved artifacts, and agent graph give a reviewer material to inspect. Its dashboard stays local by default and uses a tokenized session link. Binding the viewer to all interfaces is possible, but the README warns that the link grants access to scan data and should be protected by a firewall.

The weak point is knowing what the agents did not complete. Issue #1155 describes a specialist agent refusing its task while Strix treated the response as ordinary text and continued. The reporter's concern is false confidence: a report may appear complete even though a specialist never tested its assigned area. Until that path is fixed and covered by tests, require explicit task status in any review process and treat missing evidence as missing coverage.

Model choice adds another variable. Different providers may refuse offensive instructions, use tools differently, or spend very different amounts during a long multi-agent run. Strix exposes reasoning settings and supports several providers, but provider choice is part of the security methodology. Record it alongside the target, commit, scope, and result.

Project health and the decision

The repository was pushed on August 27, 2026, and release v1.5.3 arrived on August 10. GitHub showed 318 open issues and pull requests combined, with fresh issue and pull-request activity on August 27. The project is moving quickly, which helps fixes land but calls for pinned versions and a repeatable acceptance suite.

Documentation covers the CLI, viewer, API targets, CI, provider settings, MCP connections, agent skills, and the legal boundary around authorized testing. It is good enough to start a lab trial. The failed pinned suite keeps us from calling the source checkout dependable without qualification.

Use Strix as a second set of hands for an experienced security reviewer. Keep it isolated, verify every claimed exploit, and audit whether each assigned test finished. Teams wanting predictable signature or template checks should start with Nuclei or ZAP; teams prepared to supervise an offensive agent will find more ambitious coverage here.

Alternatives

ProjectWhat it isPick it when
OWASP ZAPA mature web application scanner with proxy, automation, and active-scan modes.pick this instead when you want deterministic web scanning without paying for an LLM on every run.
Nuclei gh↗A template-driven scanner for known exposures, misconfigurations, and vulnerability checks.pick this instead when repeatable, reviewable templates matter more than agent-written exploit paths.
PentAGI gh↗An agent-based penetration-testing system with a web interface and tool execution.pick this instead when you want to compare another self-hosted multi-agent approach and are prepared for a larger service stack.

What people are saying

  1. [github-trending] usestrix/strix

Sources

  1. Strix README
  2. Strix v1.5.3 release
  3. Custom LLM API base failure report
  4. Specialist refusal handling report

More ai tools reviews

iFixAi · dream-loop · Codex-Minecraft-Gameplay · kun · screenwriting-skills · holo-card-studio · the whole board →