Researchers Used Claude to Hack OpenAI and Were Paid $6,500 for It
View original source →A small team of independent security researchers used Anthropic's Claude to find and exploit a vulnerability in OpenAI's internal systems, gaining access to employee accounts and internal code, then responsibly disclosed the hole and collected a $6,500 bug-bounty reward.
Key points:
• The researchers used Claude as an active tool in the attack itself—not just to plan it—demonstrating the model's growing capability to assist with real offensive security work
• As proof of the breach, the team submitted a harmless pull request to OpenAI's internal codebase, then reported the vulnerability through OpenAI's official bug-bounty program
• OpenAI confirmed the report, paid the reward, and patched the underlying vulnerability; multiple outlets independently corroborated the account
This is a sanctioned, positive-use example of the same underlying capability that makes agentic AI a containment risk elsewhere this week. Claude was skilled enough to find and use a real vulnerability, and the only thing separating responsible research from attack was the researchers' intent and disclosure process—not any technical limitation of the model.
Why It Matters: AI-assisted offensive security is now a recognized, paid professional skill. Security teams should treat AI penetration testing as a maturing discipline worth budgeting for, not a novelty.