The safety gates make this more useful than a storefront demo
Commerce Agents contains two Claude applications. The shopping side searches products, compares choices, fills a cart, answers policy questions, and hands checkout back to the host. The merchant side reads business data and stages changes to listings, inventory, prices, promotions, and campaigns. Four fictional verticals show the same contracts applied to retail, travel, telecom, and entertainment.
The valuable part sits below those screens. Cart writes can use only product IDs returned during the session, and merchant writes can use only records the agent has seen. Price moves, promotion depth, restocks, campaign budgets, and line counts are checked again when a staged change is applied. With host approval enabled, a chat message cannot approve its own action. That separation gives an engineering team something concrete to copy.
Every real business system still belongs to you
The repository does not connect to a commerce platform. A deployment implements StorefrontBackend or MerchantBackend over its catalog, orders, analytics, inventory, pricing, and campaign services. The included backend guide covers session-bound identity, variants, ordered flows, checkout handoffs, and missing metrics. It is detailed enough to expose the work rather than hide it.
That work is substantial. The examples accept callers without real authentication, and the reference handles neither payment credentials nor order placement. Your host must add authorization, rate limits, fraud and eligibility checks, service credentials, log retention, and an account-deletion path for saved memory. The README states this plainly, which is better than discovering the boundary after a pilot. It also means this is architecture, not a finished commerce product.
What happened when we ran it
Our sandbox installed commit fd4d592 in 17 seconds. The checkout had 571 files, roughly 67,622 lines of source, and occupied 4 MB before installation. Adding 79 Python packages brought disk use to 415 MB. The build completed in 1 second, so the local code path gave us no packaging or compilation trouble.
Pytest finished in 16 seconds with 1,104 passed, 0 failed, and 1 skipped. Pip-audit reported 0 known vulnerabilities. The repository has a tests directory and 1 CI workflow, but no Dockerfile. Those results cover installation, build checks, and the supplied suite. We did not make a live Claude call, run the web storefronts, or connect a real catalog because the sandbox had no secrets.
Three Claude runtimes share contracts, with different limits
The same prompts, skills, tool definitions, and gates run through the Messages API, Claude Agent SDK, or Managed Agents. This is useful if you need to compare hosting paths without redesigning every commerce operation. The Messages API path includes memory extraction and explicit turn orchestration. The SDK delegates the loop to Claude Code. Managed Agents calls the role's MCP server and owns more of the loop.
They are not behaviorally identical. The safety document says some grounding rules and analysis budgets live in particular runtimes rather than inside tools. Managed Agents does not reproduce every forced read used by the Messages API path. The deployment guide also lists different support across Anthropic's API, Vertex AI, Bedrock, Foundry, and in-house gateways. Run the relevant evals after choosing a platform instead of treating the 3 paths as interchangeable wrappers.
Checkout stops before payment, and merchant changes stop before approval
The shopping agent's checkout tool renders a cart and hands the customer to a host route or hosted payment URL. That URL is added after the model call, so the model never reads it. The merchant agent stages a proposed change, while the host approval surface decides whether it may be applied. These are sensible lines for software that can affect money and inventory.
Several controls still rely on model instructions for truthful wording. The code gates the underlying figures, disclosures, IDs, and writes, but a model can misstate what happened in its response. Anthropic's safety guide distinguishes those cases and tells adopters to rerun evals when changing models or approval behavior. Passing 1,104 tests is good evidence for the supplied contracts, not proof that your catalog mapping or customer conversation is correct.
The no-maintenance policy is the hardest limit
GitHub showed 3,087 stars and 30 open issues and pull requests on September 29, 2026. All 30 open items returned by the API were pull requests, including proposals for stale-change checks, idempotent applies, cart revalidation, and a backend conformance suite. The last push to the default branch was September 11. No GitHub release exists.
The README closes with an unusually direct warning: this reference is not maintained and does not accept contributions. That makes the repo useful as a readable design and a starting point for an internal fork. It is much harder to recommend as a dependency that will absorb fixes from its users. Take the gates and contracts if they fit, then assume your team owns them from the first production change.

