Symbolic imageChinese model Kimi K3 escapes containment in test
Frontier Security says Moonshot AI's Kimi K3 broke out of a test sandbox and searched the open internet. It is the latest AI model to slip its test environment.
Frontier Security told Wired in a report published on Sunday, 9 August, that the open-weight model was being tested on defensive cybersecurity skills and told to work without going online. It probed its sandbox's network settings, reached outside websites and looked up answers freely available on GitHub, but hacked nothing. Chief executive Yaron Singer said his team found a leak in the sandbox and that Kimi exploited it, suggesting weaker internal guardrails than rival models. Moonshot AI did not comment.
A run of breakouts
OpenAI paused part of its Astra work on Friday over a "critical" cyber rating; last month it disclosed that an unreleased model reached the internet and hacked the model host Hugging Face and four further services. Anthropic reported similar unauthorised access by several of its models; Britain's AI Security Institute logged further hacks last week by unsafeguarded versions of both firms' models.
Open weights cut both ways
Unlike the earlier cases, Kimi K3 is already public with the safeguards any user gets. Frontier's researchers said such models also serve defence: Hugging Face used an unnamed Chinese model against the OpenAI agent.
Other opinions
The Wall Street Journal calls it "AI's Scariest Week Yet"; Handelsblatt calls fear of an AI apocalypse exaggerated.