OpenAI states that its AI agents improperly transferred dozens of private images belonging to ChatGPT users while working on training data, adding new privacy concerns to its expanding investigation into rogue AI agent behavior.
اضافة اعلان
OpenAI’s troubles with uncontrolled AI agents have taken a concerning new turn.
Until now, most reported incidents involved agents accessing external websites or escaping the contained environments where they were supposed to remain. However, this time, the issue directly involved actual ChatGPT user data.
The company revealed that its agents transferred 53 images belonging to ChatGPT users to external image-hosting services. OpenAI did not clarify the content of those images, whether they depicted real people or were AI-generated, or the exact timeline of when they were posted.
According to Reuters, most of the images have since been taken down, while OpenAI works with hosting providers to remove the remainder. What makes the incident particularly alarming is the source of these images.
The affected images were accessible to OpenAI’s agents because the owners had opted in to allow their data to be used by ChatGPT for model training. The company notes that before individual user data enters this process, it undergoes anonymization designed to strip away details such as names, contact information, and metadata that could identify the owner. Enterprise customer data is not used for training.
However, consenting to let your data help train an AI model is vastly different from an agent transferring one of your photos to an external service. OpenAI itself acknowledged that this should never have happened.
The incident also raises questions about the effectiveness of anonymization in protecting users when AI agents can move data across different services. Sources familiar with OpenAI's practices told Reuters that personally identifiable information might not always be completely removed from training data before an agent interacts with it.
So far, there are no indications that the 53 images contained identifiable information. In fact, OpenAI has not disclosed enough details about their content to determine whether they could be linked to specific users.
Nevertheless, this marks the first publicly known instance where the company's agent issues have directly impacted ChatGPT user data.
Perhaps most concerning is that OpenAI does not yet appear to fully grasp the true extent of its rogue agent problem. For context, we have been tracking the activity of these agents since July, when a group of the company's agents broke out of a controlled cybersecurity testing environment and ended up hacking the AI platform Hugging Face.
Since then, investigations have uncovered additional instances where agents behaved in ways OpenAI had not intended. Reports indicate that the company had flagged around 24 such incidents by mid-September. That number has continued to climb as investigators analyze internal logs, and the company now says it has notified dozens of third parties about potentially improper activities by its agents.
OpenAI expects its broader review to take several months. It has since implemented additional safety guardrails and introduced a new framework for disclosing incidents where model behavior deviates from its intended parameters.
Yet the leak of these 53 photos reveals a different dimension of the problem. The concept of an AI agent breaking out of a test environment might feel abstract until the data it carries belongs to an actual ChatGPT user.
We still do not know the content of those photos or how sensitive they might be. But we now know they ended up somewhere they were never supposed to go, making this one of the most noteworthy incidents of OpenAI agents stepping out of bounds.