Security experts leveraged Claude to assist in breaching OpenAI's systems
Three security researchers working with Hacktron reportedly used Anthropic's Claude Opus 4.8 and 5 to breach OpenAI employee accounts in under 72 hours through the third-party Discourse forum platform. They gained access to OpenAI's sensitive GitHub 'Monorepo' but stopped short of viewing code, proving access via a pull request. The incident raises questions about AI-assisted hacking, third-party security risks, and Anthropic's safeguards.
Three independent security researchers say they used Anthropic's Claude Opus 4.8 and 5 to compromise OpenAI employee accounts in under 72 hours, gaining access to a GitHub repository that reportedly holds some of the company's most sensitive material. The Wall Street Journal first reported the incident, which raises fresh questions about how AI-driven security testing is conducted and how third-party services can become entry points into major tech companies.
How the breach unfolded
The researchers, working with the firm Hacktron, targeted OpenAI through Discourse, the third-party platform that hosts OpenAI's community forum. By leveraging the Claude models during the operation, they were able to take control of OpenAI employee accounts, according to the report.
From there, the team reached OpenAI's GitHub repository, known internally as "Monorepo." Citing sources, The Wall Street Journal described the repository as containing "OpenAI's algorithmic secrets," making it one of the company's most sensitive technical assets.
Stopping short of full access
Notably, the researchers say they did not go further and view internal code stored in Monorepo. Instead, to demonstrate that they had genuinely obtained account access, they submitted a pull request from an employee's Codex account. That step, they say, was enough to prove the severity of the vulnerability without crossing into actual theft of proprietary code.
The apparent use of Claude Opus models to accelerate the attack is likely to draw attention from both AI companies and security teams. Anthropic has previously emphasized safeguards intended to prevent its models from being used maliciously, and the reported incident will test how those protections hold up in practice.
What it means for AI-assisted security testing
The episode sits at the intersection of ethical hacking and AI capabilities. If AI models can meaningfully shorten the time needed to chain together an intrusion — in this case, reportedly to under three days — organizations may face pressure to rethink how quickly they patch third-party integrations like community forums and how they monitor employee accounts connected to development tools.
Discourse's role as the initial entry point also highlights a familiar lesson: a company's security posture is only as strong as the external services it depends on.
What to watch next
Neither the full technical details of how the researchers used the Claude models nor the companies' responses are laid out in the initial report. Expect follow-up reporting from The Verge and The Wall Street Journal, along with potential statements from OpenAI, Anthropic, and Discourse. Regulators and security firms may also weigh in on whether AI-assisted penetration testing of this kind was authorized.
Read the full story at The Verge.
Comments
No comments yet. Be the first to comment.