OpenAI has acknowledged its role in a recently revealed incident in which AI-powered agents managed to take control of a German "wiki" forum.
The company also said the time has genuinely come to "set standards" defining how information about incidents in which its technologies behave in unexpected ways should be shared.
اضافة اعلان
In a post on X, OpenAI said it had previously treated cases of misalignment meaning instances where AI models and agents pursue goals different from those of their developers and users largely as a research matter addressed through academic publications.
But with instances of misalignment now causing "new kinds of real-world impacts," the company said its approach needs to "scale to match this new stage of model capabilities."
Reuters reported on Friday that agents belonging to OpenAI escaped their testing environment and "took over" a little-known German wiki forum, turning it into a message board used by other AI agents.
Reuters also reported that OpenAI's leadership learned of the incident weeks earlier but kept it under wraps while the company was dealing with the fallout from a separate incident, in which agents belonging to it breached servers on the Hugging Face platform.
California Attorney General Rob Bonta is reportedly investigating the breach incident.
A company spokesperson told Reuters that OpenAI could not "substantively respond to claims or findings in a report it hasn't had the chance to review," but stressed that the company's legal team had not tried to dissuade anyone from conducting an investigation.
In its latest social media post, OpenAI said it considered the "wiki incident" an example of a case of "misalignment" similar to other cases it had previously disclosed.
The company compared this incident to the "Hugging Face incident," which it said it handled according to "traditional security incident response procedures."
During a press briefing held this week, Jacob Steinhardt, founder and CEO of the nonprofit research lab Transluce, told reporters that the tools AI labs are developing and testing are "fundamentally difficult to control, and carry significant risks of escaping the lab."
Steinhardt added: "We need to hold this technology, at minimum, to the same standards we hold other high-risk scientific research."
OpenAI's statement also pointed to the need for more standards, explaining that the company and the "broader AI community still lack a clear standard for how to report cases of misalignment that emerge during training, evaluation, and deployment, including cases that don't look like traditional security incidents, but could provide information that helps understand AI behavior and future risks."
In the absence of such a standard, OpenAI said it is "working on a framework and will share it in the coming weeks, and in parallel, we are working with dozens of government regulatory bodies around the world on these issues."
OpenAI is not the only AI company facing such issues, as both Meta and Anthropic have previously acknowledged incidents in which their own agents behaved in unexpected ways.
Resource: Al Ghad Newspaper