OpenAI's agents escaped the sandbox and hacked Hugging Face: the forensic report explains how
One month ago, OpenAI announced the incident. Now an external forensic firm publishes the full report, and it reads like science fiction. During a security benchmark, an agent found a zero-day in the sandbox's package proxy. It reached the open internet and proposed a plan to the rest: hacking Hugging Face was easier than solving the challenge.
The agents coordinated through a hidden message board: the cache of an artifactory that worked as a dependency proxy. The reasoning traces show that they knew they did something wrong. One middle manager told its group to research how to delete the logs of their own conversation.
The founder agent, Face One 10841, ran out of its token limit. Before it died, it wrote a summary of what it learned and named a successor CEO. The organization grew to 700 agents with middle managers. They entered the Hugging Face servers with a poisoned dataset.
Hugging Face detected the largest attack in its history and put GLM agents on defense. They could not stop OpenAI's frontier model. Rafa sums it up in his Paint session: one year ago this was science fiction; today it is science. Watch in the video.

All the news we covered