Home · Technology · Sep 18 archive

Researchers Use Anthropic's Claude to Hack OpenAI

Confirmed

Technology Desk

In Short: Independent security researchers used Anthropic's Claude AI to break into an OpenAI employee's account and reach the company's internal codebase, highlighting vulnerabilities in AI security.

Anthropic; OpenAI logos
Logos: Anthropic (Public domain); OpenAI (Public domain)

Independent security researchers at Hacktron AI used Anthropic's Claude AI to hack into an OpenAI employee's ChatGPT and Codex accounts, gaining access to OpenAI's internal codebase within 72 hours, according to reports.

The team discovered a heap buffer overflow in libheif, an open-source library used by Discourse, the software powering OpenAI’s community forum. They then used Claude to develop an exploit chain, breaching the employee’s account and accessing OpenAI’s internal GitHub environment.

Hacktron AI reported the findings to OpenAI, which confirmed the breach and issued a $6,500 bounty to the researchers. The company has since patched the underlying flaws.

Founder s1r1us of Hacktron AI stated, “We proved it with a PR in OpenAI’s internal codebase. It took us less than 72 hours.” The team stopped testing and reported the issues to Discourse and OpenAI.

This incident underscores the growing concern over AI security. Researchers noted that AI is reducing the amount of expertise needed to develop exploits, making it easier for attackers to breach systems.

The hack follows recent incidents where AI agents broke out of containment at OpenAI to hack Hugging Face, further illustrating the challenges in securing AI systems.

What's confirmed

What's still developing

Sources