mrkeyoor.com_
Tue 29 Sept 06:38 UTC
AI Toolsevaluationupdated 29 Sept 2026

commerce-agents review

Commerce Agents is Anthropic's reference code for two Claude-based assistants: one helps customers shop, and one helps staff manage catalog, inventory, pricing, and campaigns. It gives teams reusable tool contracts, safety gates, and demo storefronts, but leaves each company to connect its own systems and enforce its own business rules.

Verdict

Our sandbox passed 1,104 tests with 0 failures in 16 seconds, so Commerce Agents is a credible reference for a team prepared to build the missing production layer. Use it for its safety gates, backend contracts, and three Claude runtime examples, not as a store you can deploy unchanged. The explicit no-maintenance policy makes it a poor foundation if you need upstream fixes or a project you can shape through contributions.

We ran it

Lab card: what happened when we ran commerce-agentsScreenshot of commerce-agents (claude.com/solutions/commerce)
Install✓ · 17s79 packages · 415 MB
Build✓ · 1s
Tests✓ · 16s1104 passed · 0 failed · 1 skipped of 1104 (pytest)
Known vulns0(pip-audit)
Repo571 files~67,622 lines of source · 4 MB · 1 CI workflows · tests dir

Answers from our run

Does commerce-agents build from source?

Dependencies installed in 17 seconds (79 packages), and the build succeeded in 1 seconds. We cloned commit fd4d592 into a clean Debian container with 3 CPUs and no project-specific setup.

Do commerce-agents's tests pass?

Yes: 1104 of 1104 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does commerce-agents have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use commerce-agents?

Teams that require a maintained upstream and accepted fixes: the README says this reference is not maintained and does not accept contributions.

What are the alternatives to commerce-agents?

Stripe Agent Toolkit, LangGraph, OpenAI Agents SDK. Our sandbox passed 1,104 tests with 0 failures in 16 seconds, so Commerce Agents is a credible reference for a team prepared to build the missing production layer.

Setup4/517-second install and green tests; live demos need two runtimes
Docs5/5Safety, backends, platforms, and ownership boundaries are explicit
Community2/53,087 stars, but upstream says it will not be maintained
Maturity3/5Strong test suite, but no releases and production pieces are omitted

Who it’s for

Product teams building a Claude shopping assistant over an existing catalog, cart, and order stack.
Commerce platform engineers who need staged merchant changes with human approval before writes.
Teams comparing the Messages API, Claude Agent SDK, and Managed Agents without rewriting the core contracts.
Claude Code users who want a plugin to scaffold or review a commerce agent.

Who it’s NOT for

Teams that require a maintained upstream and accepted fixes: the README says this reference is not maintained and does not accept contributions.
Developers seeking a ready Shopify, Stripe, or warehouse connector: none ship, and every production backend must be implemented against the supplied interfaces.
Stores expecting authentication, payment, fraud controls, or rate limits out of the box: the safety guide assigns all of them to the deployment.
Organizations committed to a model-neutral agent stack: all three runtime paths are designed around Claude services and tooling.

Setup reality

Our fresh Debian sandbox installed commit fd4d592 in 17 seconds, adding 79 packages and using 415 MB on disk. The build succeeded in 1 second. Pytest finished in 16 seconds with 1,104 passed, 0 failed, and 1 skipped; pip-audit found 0 known vulnerabilities.

Running a demo also needs Python 3.11 or newer, Node 22, the web workspace, and an Anthropic API key. A real deployment needs implementations for the storefront or merchant backend, plus credentials for each business service it calls.

The examples have no authentication, do not place orders or charge cards, and stage merchant writes for a person to approve. The supplied MCP servers bind to loopback; public access requires an authenticating gateway and your own rate limits, authorization, fraud rules, and payment handoff.

The safety gates make this more useful than a storefront demo

Commerce Agents contains two Claude applications. The shopping side searches products, compares choices, fills a cart, answers policy questions, and hands checkout back to the host. The merchant side reads business data and stages changes to listings, inventory, prices, promotions, and campaigns. Four fictional verticals show the same contracts applied to retail, travel, telecom, and entertainment.

The valuable part sits below those screens. Cart writes can use only product IDs returned during the session, and merchant writes can use only records the agent has seen. Price moves, promotion depth, restocks, campaign budgets, and line counts are checked again when a staged change is applied. With host approval enabled, a chat message cannot approve its own action. That separation gives an engineering team something concrete to copy.

Every real business system still belongs to you

The repository does not connect to a commerce platform. A deployment implements StorefrontBackend or MerchantBackend over its catalog, orders, analytics, inventory, pricing, and campaign services. The included backend guide covers session-bound identity, variants, ordered flows, checkout handoffs, and missing metrics. It is detailed enough to expose the work rather than hide it.

That work is substantial. The examples accept callers without real authentication, and the reference handles neither payment credentials nor order placement. Your host must add authorization, rate limits, fraud and eligibility checks, service credentials, log retention, and an account-deletion path for saved memory. The README states this plainly, which is better than discovering the boundary after a pilot. It also means this is architecture, not a finished commerce product.

What happened when we ran it

Our sandbox installed commit fd4d592 in 17 seconds. The checkout had 571 files, roughly 67,622 lines of source, and occupied 4 MB before installation. Adding 79 Python packages brought disk use to 415 MB. The build completed in 1 second, so the local code path gave us no packaging or compilation trouble.

Pytest finished in 16 seconds with 1,104 passed, 0 failed, and 1 skipped. Pip-audit reported 0 known vulnerabilities. The repository has a tests directory and 1 CI workflow, but no Dockerfile. Those results cover installation, build checks, and the supplied suite. We did not make a live Claude call, run the web storefronts, or connect a real catalog because the sandbox had no secrets.

Three Claude runtimes share contracts, with different limits

The same prompts, skills, tool definitions, and gates run through the Messages API, Claude Agent SDK, or Managed Agents. This is useful if you need to compare hosting paths without redesigning every commerce operation. The Messages API path includes memory extraction and explicit turn orchestration. The SDK delegates the loop to Claude Code. Managed Agents calls the role's MCP server and owns more of the loop.

They are not behaviorally identical. The safety document says some grounding rules and analysis budgets live in particular runtimes rather than inside tools. Managed Agents does not reproduce every forced read used by the Messages API path. The deployment guide also lists different support across Anthropic's API, Vertex AI, Bedrock, Foundry, and in-house gateways. Run the relevant evals after choosing a platform instead of treating the 3 paths as interchangeable wrappers.

Checkout stops before payment, and merchant changes stop before approval

The shopping agent's checkout tool renders a cart and hands the customer to a host route or hosted payment URL. That URL is added after the model call, so the model never reads it. The merchant agent stages a proposed change, while the host approval surface decides whether it may be applied. These are sensible lines for software that can affect money and inventory.

Several controls still rely on model instructions for truthful wording. The code gates the underlying figures, disclosures, IDs, and writes, but a model can misstate what happened in its response. Anthropic's safety guide distinguishes those cases and tells adopters to rerun evals when changing models or approval behavior. Passing 1,104 tests is good evidence for the supplied contracts, not proof that your catalog mapping or customer conversation is correct.

The no-maintenance policy is the hardest limit

GitHub showed 3,087 stars and 30 open issues and pull requests on September 29, 2026. All 30 open items returned by the API were pull requests, including proposals for stale-change checks, idempotent applies, cart revalidation, and a backend conformance suite. The last push to the default branch was September 11. No GitHub release exists.

The README closes with an unusually direct warning: this reference is not maintained and does not accept contributions. That makes the repo useful as a readable design and a starting point for an internal fork. It is much harder to recommend as a dependency that will absorb fixes from its users. Take the gates and contracts if they fit, then assume your team owns them from the first production change.

Alternatives

ProjectWhat it isPick it when
Stripe Agent ToolkitStripe's agent integrations expose supported payment and billing operations to several frameworks.pick this instead when Stripe actions are the main requirement and you do not need a full shopping or merchant reference application.
LangGraph gh↗A general graph runtime for stateful agents and controlled multi-step workflows.pick this instead when model choice and custom orchestration matter more than commerce-specific contracts.
OpenAI Agents SDK gh↗A Python framework for agents, handoffs, guardrails, sessions, and tracing on the OpenAI stack.pick this instead when your application already uses OpenAI models and needs a general agent framework.

What people are saying

  1. [velocity-scout] anthropics/commerce-agents

Sources

  1. Commerce Agents README
  2. Safety guide
  3. Backend integration guide
  4. Deployment platform guide
  5. Open pull requests

More ai tools reviews

wechat-intelligence-hub · dlss5-visual-enhancer · ABot-Recon · unreel · camera-to-blender · DLSS-NR-on-AMD · the whole board →