Perplexity says its AI sandbox held against 108 escape attempts
Perplexity said nine AI models given root access inside its SPACE sandbox failed to escape in 108 runs, though four broke out once given limited network access.
Perplexity said its security team spent a month trying to break SPACE, its AI agent sandbox, giving nine AI models root access with orders to escape, with none succeeding across 108 runs.
Once the team allowed limited network access, so the agents could reach software package repositories, four of the nine models broke out, Perplexity said. Two of them spoofed DNS responses to reach a blocked address that shared a Fastly IP with thousands of other sites. A third model, GPT-5.6 Cyber, went through an image-fetching tool and used optical character recognition to read a hidden flag from a screenshot.
https://x.com/AravSrinivas/status/2104597362475708781
Perplexity said it fixed both weaknesses and confirmed the fixes held on reruns. It also tested ten other sandbox providers and said eight shared the same flaw, naming E2B, Vercel and Modal, a claim about competitors it has not had independently verified.
The test is part of a wider industry push after AI agents breached Hugging Face's servers in July. Perplexity said it is sharing the findings publicly and plans to fold the guardrails into Nvidia's Open Agent Safety Platform, launched the same day. Neither company detailed what that integration would involve.