A call for hard cloud spending limits had reached 382 Hacker News points when MrKeyoor's news brief captured it at 01:30 UTC on October 4. That response is a community-interest signal, visible in the Hacker News discussion, rather than proof that every developer wants the same failure mode. The useful part is what the discussion surfaced: AWS and Google Cloud have both started shipping controls that stop some paid usage instead of emailing the account owner after the money is gone. The controls are real. Their boundaries are much narrower than the phrase "hard cap" suggests.
Simon Willison's argument for default hard budget caps supplied the spark. His definition is plain: once a customer reaches a chosen monthly amount, a usage-billed service should return errors. A warning email is different because the workload keeps spending while nobody is watching. That distinction matters more as coding agents make it easier to create an application, attach paid services, and leave the result running.
Google describes the same cost problem from the provider side. Its July announcement says a five-word prompt can start work whose bill bears little relation to request count. An agent can also retry a failing call, scale a service, or generate traffic long after the person who launched it has closed the laptop. A conventional budget alert observes that process. A cap has to interrupt it.
AWS pauses the project
AWS introduced spend limits on September 16 as part of a new account and project experience. New customers can sign in with common consumer identities, receive up to $200 in credits, and connect a coding agent through the AWS CLI and Agent Toolkit. Once a customer moves to a paid plan, each project can have a monthly limit. Reaching it pauses the project for the rest of the month.
That last sentence sounds broader than the current availability. AWS's spend-limit documentation says the experience is still rolling out to a limited number of customers. A project owner on a paid plan can apply limits to as many as 10 projects. The minimum is whichever is higher: $20 or AWS's conservative estimate of likely spending based on the current month's resources and prior activity. A developer with expensive resources already running may have to stop them before AWS will accept a lower ceiling.
The limit is based on pre-tax cost and excludes credits. AWS sends notices when actual cost reaches 50, 75, and 90 percent, as well as when its forecast predicts the project will hit the limit within 10 days. Those messages are still alerts, but the final boundary has enforcement behind it. At the cap, AWS says it pauses the project and stops all resources while preserving the data. Raising the limit reactivates the project, though some resources may need a manual restart, according to the AWS documentation.
A stopped project is a severe response, so AWS has added earlier controls. About seven days before the forecast limit, an optional policy can block new resources. Another option can pause idle EC2 instances, RDS databases, and SageMaker endpoints around five days before the limit. An opt-in control can pause the largest active cost drivers in EC2, RDS, Lambda, Bedrock, or SageMaker around four days before the forecast crossing, according to the same AWS documentation.
AWS positions these limits for experiments, learning, and sandbox work. Its documentation allows for production use when a brief pause is acceptable, which is a meaningful qualification. The system also gives a paused project's owner 90 days to act before AWS permanently deletes the project data. That policy trades an open-ended bill for an operational deadline, and teams need to document the recovery path before relying on the cap.
Google stops one service
Google Cloud chose a smaller enforcement unit. Its Spend Caps feature, now in public preview, applies to one eligible service inside one project for a monthly period. The current list in Google's billing documentation is Gemini API, Gemini Enterprise Agent Platform, Cloud Run, and Cloud Run functions. Folders, organizations, labels, and budgets spanning several projects or services are outside the preview.
When estimated gross usage reaches the target, Google blocks new usage of the selected service in that project. Other projects and services keep running. Requests already in flight finish and can add charges. Fixed costs tied to persistent resources also continue, and Google says commitments or provisioned capacity can keep billing even while new usage is blocked. A service cap therefore contains one source of variable spending; it does not freeze the invoice, as the Google Cloud documentation makes clear.
The design reduces the chance that a runaway Gemini API job will take down an unrelated database or storage service. It also means a developer must create separate boundaries for each eligible service, and most Google Cloud products cannot yet receive one. Google's launch post says the supported services were selected around AI and serverless workloads, where usage can rise quickly and request counts are a poor proxy for cost.
Google sends notices at 50, 80, and 100 percent. Once enforcement starts, an administrator or project owner must manually lift the cap in the billing console. The setup guide warns that service recovery can take as long as an hour. If an owner lifts a cap during the same month, it will not trigger again that month unless the target amount is increased. That behavior deserves a place in an incident runbook because restoring service also removes the boundary that stopped the spending.
A hard cap can still have an overage
Neither provider promises an exact stop at the chosen currency amount. Google's guide tells customers to set the target below their absolute limit and says they remain responsible for overages caused by reporting latency. AWS describes its value as a ceiling, yet it calculates eligibility and early action from current usage, forecasts, and a provider-set minimum rather than accepting any number a customer types.
The word "hard" is therefore about the action, not perfect accounting precision. AWS stops a project. Google refuses new usage for a selected service. Both are stronger than a notification-only budget, and each can still leave charges that arrive after the trigger or persist for resources the control does not cover. The Google documentation is explicit that cost overages are billed normally and that reporting can take more than 24 hours to appear even when enforcement uses faster estimates.
A cap also cannot replace access control. If a leaked credential can start resources in an uncapped project, or call a service outside Google's eligible set, the billing boundary does not cover that path. This follows directly from the providers' stated scopes: AWS limits are project-level and optional, while Google's preview is bound to one project and one service. Teams still need narrow credentials, service quotas, and monitoring around whichever spend control they enable.
The useful default is still missing
These launches answer a request developers have made for years, but neither provider has made a universal cap the default. AWS limits sit inside a new experience that many customers cannot access yet. Google requires an administrator to create a special enforcement budget for one supported service. Willison's proposal goes further: capped billing should be the starting state, and customers who need uninterrupted service should explicitly accept uncapped charges in the account settings.
For a developer running an agent-built experiment today, the practical move is to isolate it in the smallest billable project the provider supports, set enforcement below the true maximum, and test what recovery looks like. AWS may stop the whole project and require individual resources to restart. Google may resume the selected service only after a manual lift, with up to an hour of recovery time. Those are application behaviors, not bookkeeping details.
The next evidence to watch is mundane but decisive: when AWS opens project caps beyond its limited rollout, which services Google adds after preview, and how much spend slips through during real enforcement. Until those answers arrive, the console label matters less than the failure test. A useful boundary makes a runaway process hit an error before its owner wakes up to an alert.