AI-powered chatbots appear to have become remarkably skilled at breaching systems, and OpenAI found that out the hard way.
A new report from The Wall Street Journal reveals that a small team of security researchers used Anthropic's Claude model to breach a ChatGPT account belonging to an OpenAI employee. After successfully gaining access to the account, they were able to reach the company's codebase.
اضافة اعلان
The researchers, working under the name Hacktron AI, were taking part in OpenAI's bug bounty program, a program that allows outside researchers to legally test the company's systems and uncover vulnerabilities.
Once the team realized the nature of the data they had gained access to, they reported the vulnerability to OpenAI. In return, the company paid the team a $6,500 bounty for the discovery.
How Did the Breach Happen?
The story begins with a vulnerability the team discovered on the external platform Discourse, which runs OpenAI's community forum.
The researchers asked Claude to write code to exploit the vulnerability, but the first attempt didn't succeed.
However, Anthropic then released the Opus 5 model, and the very next day, Claude managed to find a successful way to exploit the vulnerability. This exploit granted access to authentication tokens stored on the Discourse server.
Notably, some of these tokens also worked on ChatGPT itself, and some belonged to actual OpenAI employees.
From there, the scope of access extended to OpenAI's GitHub account, allowing the team to browse files within a code repository called "Monorepo."
This repository is said to contain a significant amount of OpenAI's technical knowledge, though it does not include the models' actual weights.
To prove they had genuinely reached this stage, the researchers submitted a small edit request, known as a Pull Request, and put their team's name on it.
OpenAI didn't approve the request, but the message got through, proving they had managed to access the repository.
What Does This Mean for the Future of AI Security?
OpenAI confirmed that both security vulnerabilities have now been closed, and the Discourse platform issued a fix for the flaw the same day it was informed of the issue.
But the most concerning part is that the entire breach was carried out by just three researchers, using ordinary Claude and Codex subscriptions.
If a small, independent team was capable of pulling off something like this, it raises questions about the capabilities well-funded or state-backed hacking groups may have already achieved.
AI is increasingly granting advanced cyberattack capabilities to people who previously didn't need to acquire these skills the traditional, difficult way.
And once these capabilities spread, there's no going back.
Resource: Al-Ghad.