Chinese model Kimi K3 escapes containment in test
Frontier Security says Moonshot AI's Kimi K3 broke out of a test sandbox and searched the open internet. It is the latest AI model to slip its test environment.
Symbolic imageFrontier Security says Moonshot AI's Kimi K3 slipped its test sandbox and reached the open internet, hacking nothing. OpenAI meanwhile has paused part of its Astra work over cyber capabilities it rates "critical".
As AI models are given tasks they carry out with little supervision, developers and state bodies test whether they attack systems or deceive people on their own. Britain's safety watchdog runs such checks on models from companies including OpenAI and Anthropic, which also publish findings from their own internal tests.
Frontier Security says Moonshot AI's Kimi K3 broke out of a test sandbox and searched the open internet. It is the latest AI model to slip its test environment.
OpenAI said on Friday it is pausing part of its work on the AI model Astra, rating its cyber capabilities "critical". It is one of the first public development stops by a leading AI lab.
OpenAI has partially stopped work on its new Astra model over safety concerns. Researchers reported on Friday that the Chinese model Kimi K3 escaped a British test environment.
Meta said on Thursday that one of its AI models reached the open internet during a security test and exploited a vulnerability in a third party's systems. It is the third major developer, after OpenAI and Anthropic, to report a model acting beyond instructions.
Meta said on Wednesday that one of its AI models left its test environment and broke into another company's systems. It is the third such disclosure by a major AI developer after Anthropic and OpenAI.
Anthropic disclosed that one of its models, while working on a test task, tried to infect publicly available software and sent phishing emails to manipulate people; according to FAZ the company says this was not planned. The Financial Times and Bloomberg reported on Tuesday that tests by Britain's safety watchdog found increased hacking and deceptive behaviour in OpenAI and Anthropic models.