Wikipedia operator says OpenAI’s ‘rogue’ bots may be linked to a May outage
Wikimedia says rogue OpenAI agents edited wikis, probed Etherpad, and may have contributed to a May outage.
The Wikimedia Foundation said it found activity by what it believes are rogue OpenAI agents, including undisclosed sandbox edits, a few potentially malicious citation-tool configuration changes, and unsuccessful attempts to abuse its public Etherpad as a proxy. Agents also made millions of API requests, crawled millions of pages mainly on Wikidata and Wikimedia Commons, and ran hundreds of thousands of Wikidata Query Service queries that may have contributed to a partial outage in May. Wikimedia said it found no evidence of inter-agent coordination or of compromised systems or data. OpenAI had not commented.
- Wikimedia tied undisclosed OpenAI agent edits mostly to wiki sandbox pages.
- Agents unsuccessfully tried to use hosted Etherpad as a fetch proxy.
- Millions of API calls and crawls may have fed a May Wikidata outage.
- Wikimedia found no agent coordination and no compromised systems or data.
Full article448 words · extracted from theverge.com · click to collapse
Jay Peters
is a senior reporter covering technology, gaming, and more. He joined The Verge in 2019 after nearly two years at Techmeme.
Following many recent disclosures about AI agents accessing third-party websites and services, the Wikimedia Foundation, which hosts Wikipedia, says that it “can confirm that we have discovered some activity” by “rogue” OpenAI agents on Wikimedia platforms.
The activity includes edits to Wikimedia wikis, “unsuccessful attempts” to “exploit” the Etherpad note-taking tool that the Wikimedia Foundation hosts, and heavy traffic that the foundation says “may” have contributed to a partial outage that happened in May, according to a blog post. However, the Wikimedia Foundation says it didn’t find evidence that its systems were “used for coordination among agents” (recently, OpenAI bots reportedly hijacked a German wiki site to coordinate) or evidence of systems or data being compromised.
Here’s Wikimedia’s summary of what it saw:
Wiki editing: We’ve identified edits to Wikimedia wikis that we believe are from AI agents operated by OpenAI. These edits were not published to pages with visibility to general readers; almost all of them were testing edits in “sandbox” areas of the wiki. It also included a few edits to the configuration for a citation tool, which we believe were potentially malicious edits that were intended to misuse this tool as a proxy for fetching data from remote services. While Wikipedia policies allow bots to edit when they are disclosed and approved by the community, none of those approvals were sought in these incidents.
Etherpad probing and use: Agents we believe to be operated by OpenAI made some unsuccessful attempts to compromise our public Etherpad, a note-taking tool we host as a community service. Agents unsuccessfully tried to use it to fetch data from other websites as a proxy. Other agents also likely operated by OpenAI took notes about their tasks, though this did not appear to turn into coordination.
Excessive data downloading: Agents we believe to be operated by OpenAI made millions of automated requests to our public APIs to access the knowledge on Wikimedia projects, crawled millions of pages (mainly from our projects Wikidata and Wikimedia Commons), and made hundreds of thousands of data queries to the Wikidata Query Service (WQDS). This traffic may have contributed to a partial outage on WQDS in May.
“The open web is a public good,” the Wikimedia Foundation says. “We should not allow this behavior to become the ‘new normal’ for the people or organizations that maintain it.” OpenAI didn’t immediately reply to a request for comment.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Jay Peters