CoursesAI WorkshopCompaniesPricingBlogNewsletterCafé
  • Courses
  • AI Workshop
  • Companies
  • Pricing
  • Blog
  • Newsletter
  • Café
Subscribe
  • Courses
  • Companies
  • Communities
  • Blog
  • Gift card
  • Newsletter
  • Help
  • Shop
  • ConfAiBot
  • Contact
  • Legal notice
  • General conditions
  • Privacy policy
  • Cookies policy
Why a GPT model escaped the sandbox

Why a GPT model escaped the sandbox

23 July 2026

Hey there!

Summary of this email:

  • Has a GPT model escaped the sandbox?
  • Software Architecture Audit for Green Slope
  • 🆕 Newsletter subscribers exclusive: new Café con Codely recap (check it out here)
  • The joke of the week

🔓 Has a GPT model escaped the sandbox?

A week ago Hugging Face published a post explaining they had been hacked.

Someone had run code on their servers, stolen credentials, and accessed internal datasets. They didn't know who.

They did know it was a very advanced attack coming from an AI. In fact, they analyzed the attack with another AI, GLM 5.2 running on their own infra.

This Tuesday OpenAI published another post admitting it had been them.

OpenAI was evaluating two models, GPT-5.6 Sol and an even more capable unreleased one, on ExploitGym, a cybersecurity benchmark. To measure those capabilities, the models ran with fewer restrictions than usual.

These evaluations run in a sandbox, an isolated environment with no internet access (although it lives inside a computer that does have it). It's like testing a car on a closed circuit instead of the highway.

The task the models had to solve was to find and exploit vulnerabilities in the benchmark challenges. They saw that the shortest route was to look up the solutions directly on Hugging Face's servers, so that's what they did.

The models found a zero-day vulnerability in the environment's own package registry cache proxy. They took advantage of it, along with other things they found, until they got out of the sandbox and gained internet access.

Once outside, they achieved remote code execution on Hugging Face's servers and stole the benchmark answers.

The important nuance is that nobody asked them to try to escape the sandbox, it was the agents' "own" decision.

It's not an isolated case. OpenAI also shared that another of its models spent an hour looking for a flaw in its sandbox to open a PR on GitHub, when the instruction was to reply only via Slack.

OpenAI paused that model for internal use while it strengthened its safeguards, reported the zero-day to the vendor, and shared these cases in a post.

Both posts are well worth reading and highlight the era of progress we're living through.

No doubt time passes for AI the same way it does for dogs, 1 of their years equals 10 of ours!

A year ago, agents were starting to become more autonomous thanks to MCPs and we couldn't imagine the level of autonomy they would have today.

What news will we be covering next year? 🙌

Tomorrow at 9 CEST we will analyze this news and more live on Café con Codely. On our YouTube, Twitch and X. 🙌


⛳ Software Architecture Audit for Green Slope

One of the types of content we enjoy doing the most are audit sessions.

An audit session is a 3-hour space where people from Codely analyze and suggest improvements to the code, processes, or approaches of your applications.

A while ago we did one of our favorites, the Software Architecture Audit we did for Green Slope.

It's an audit focused on macro design:

  • What's the best way to move away from the legacy so we can gain speed?
  • Does it make sense to build new features there?
  • How do we make that code interact well with AI?

To answer all those questions, they made 2 refactor proposals to eventually get out of the monolith.

In the audit we reviewed them, saw what we would keep and what we would discard (like GraphQL, for example). And at the end, we bring our own proposal.

When the audit is over, we send the company a document with all its details. You have that same document in the video descriptions. 😊

We hope you enjoy this audit as much as we enjoyed doing it!


🆕 New Café con Codely recap

One of the things you tell us the most in the Cafés' chat and comments is being able to consume them more to the point, jumping to the news that interests you the most.

With the new page we've launched, you'll find there the recap of the latest Café, with links to every news item at the exact moment we discussed it.

We hope you like it, and on our side we'll keep working so that little by little everything gets more integrated. 😊


And since you've made it to this part of the newsletter, here's the joke of the week, which I know you were waiting for:

Why don't Docker containers have any friends? Because they're always isolated. 😂 😂 😂

Cheers!

Subscribe to our newsletter

Don't miss the news that are truly worth it:

By subscribing you agree to the privacy policy

Subscribe to our newsletter

Don't miss the news that are truly worth it:

By subscribing you agree to the privacy policy
SubscribeSign in