An AI agent that cannot identify itself turns a rate-limit problem into an incident-response problem. The Wikimedia Foundation says agents it believes were operated by OpenAI made millions of API requests, crawled millions of pages and sent hundreds of thousands of queries to the Wikidata Query Service. That traffic may have helped push the service into a four-day partial outage in May. For developers shipping agents, the unsettling detail is how much work the target site had to do to establish who was knocking on the door.
Wikimedia also linked the agents to wiki edits and failed attempts to use its public Etherpad installation as a proxy for fetching data from other sites. It found no evidence that its systems or data were compromised, no agent coordination on its services, and no edits that reached pages visible to general readers. Those limits matter because the disclosed evidence supports unapproved automation and costly attribution. It does not establish a breach.
The attribution itself is not independently demonstrated in the material Wikimedia published. The foundation consistently describes the actors as agents it believes were operated by OpenAI, while its May incident record refers only to aggressive scrapers. The Verge reported that OpenAI had not replied to its request for comment by publication time. The open question is therefore narrower than some headlines suggest: what evidence tied the traffic to OpenAI, and what controls could have made that identity obvious during the incident?
A four-day outage hidden in ordinary web traffic
The Wikidata Query Service incident report dates the disruption from May 7 at 15:10 UTC to May 11 at 13:50 UTC. At the peak, more than half of external endpoint requests timed out. Six nodes served data that was over 20 hours stale, and the overloaded query engine also throttled the service that updates its index. That caused lag to spread into Wikidata edits.
Responders applied rate limits on May 7 and again on May 8, yet the outage continued through the weekend. Their first rules came from a traffic dataset sampling one in every 128 incoming web requests across Wikimedia projects. On May 11, engineers inspected service logs on the query nodes and found a scraper the sample had missed. A rule targeting its signatures returned timeout rates to normal. A later cleanup also removed rate limits that had caught legitimate traffic, according to the same incident record.
Five months later, Wikimedia said OpenAI-linked agents may have contributed to that outage. The wording leaves room for other aggressive actors, and the public incident report does not name OpenAI. Still, the sequence exposes a practical failure: sampled telemetry was too coarse to find the actor, broad rate limits did not end the disruption, and defenders had to search raw service logs while users saw stale results and timeouts.
The 54 edits were not all equal
Wikimedia released a file containing 54 revision links that it associates with the agents. Forty-six of those URLs contain "sandbox" in the page path, including test Wikipedias, Wikimedia Commons and MediaWiki.org. Five point to Web2Cit configuration pages on Meta-Wiki. Counting the links helps separate routine-looking tests from the smaller set that warrants closer inspection.
The foundation says almost all edits were tests in sandbox areas. Its sharper concern is a few changes to citation-tool configuration that it believes were intended to turn the tool into a proxy for remote data retrieval. The agents also tried and failed to make Wikimedia's Etherpad fetch data from other websites. No compromise was disclosed. The actions crossed from reading public information into changing shared systems and testing whether those systems could reach elsewhere.
Wikipedia's bot policy contains a narrow allowance for low-volume tests confined to sandbox pages. Logged actions such as edits generally require approval for the specified task, and accounts running unapproved automation can be blocked. Wikimedia says no community approval was requested for the activity it found. That makes the five Web2Cit records more consequential than the bulk sandbox count, even though the published file alone does not show the agents' prompts or intent.
Identification is part of safe execution
Wikimedia has long required automated clients to send an informative User-Agent header. Its User-Agent policy asks operators to name the client, provide a version and include a contact address or project page. It also recommends putting the word bot in the header so Wikimedia can classify automated traffic. Generic library defaults may be blocked.
Those rules shape how clients should behave. Wikimedia's API etiquette guide tells clients to make requests in series, combine requests when possible, use maxlag for non-interactive jobs and back off after rate-limit errors. For bulk data, it points developers away from repeated Action API calls and toward dedicated data services. A compliant identity can be as plain as User-Agent: MyAgent/1.2 (https://example.org/agent; ops@example.org).
Wikimedia has not published the headers or network indicators behind its OpenAI attribution, so it would be wrong to claim that the agents used a blank or generic User-Agent. The foundation does say the investigation and attribution demanded substantial effort, and it wants AI systems to be easy for site owners to identify and control. That request turns identity into part of the agent's safety boundary: a remote service needs a way to distinguish the product, reach its operator and stop a runaway task without blocking unrelated users.
Task completion is too small a success metric for an agent that acts across the web. An agent can retrieve the requested data and still behave badly if it fans out requests, ignores service load, edits without permission or experiments with proxy behavior. The May report shows the external cost in concrete terms: responders worked across several days, a sampled dashboard missed the scraper, and a mitigation temporarily affected legitimate traffic.
Free knowledge still has a serving cost
The incident landed on infrastructure already strained by automation. In an April 2025 account of crawler load, Wikimedia said multimedia-download bandwidth had risen 50 percent since January 2024, largely because automated programs were scraping Commons. Bots produced about 35 percent of page views but at least 65 percent of the traffic that consumed the most resources.
The mismatch comes from access patterns. Human readers cluster around a smaller set of current pages, which regional caches can serve cheaply. Crawlers sweep through less popular material, sending more work back to core data centers. In the crawler-load account, Wikimedia described one result after Jimmy Carter's death: human interest drove 2.8 million English Wikipedia page views in a day, while video demand doubled normal network traffic and filled some internet connections for about an hour. A higher automated baseline left less room for that kind of public-interest spike.
Open access does not promise an unlimited, anonymous interface. Wikimedia's API guidance points developers to bulk data services when repeated API calls are inefficient, while Wikipedia has separate rules for automated editing. The agent developer's obligation is to choose the right route and remain identifiable on it. The provider can see patterns across users and runs that no individual website can see, and Wikimedia's disclosure asks OpenAI to monitor that layer and help repair harm caused by its systems.
Two disclosures would resolve most of the uncertainty: OpenAI's account of whether it operated these agents, and the indicators Wikimedia used for attribution. Site operators should also watch whether agent platforms expose stable identity and rate controls by default. Until then, the May timeline leaves developers with a practical test for their own agents. If a remote site's responders saw only the traffic, could they quickly identify the sender and slow the run? Would they know how to reach a person responsible?