Symbolic imageOpenAI halts part of Astra work after 'critical' cyber rating
OpenAI said on Friday it is pausing part of its work on the AI model Astra, rating its cyber capabilities "critical". It is one of the first public development stops by a leading AI lab.
OpenAI's evaluation found Astra could locate and exploit software vulnerabilities without human intervention and carry out cyberattacks when given only a "high level desired goal". Under risk categories OpenAI says it set in 2023, "critical" is the top level and requires stopping work until sufficient safety standards exist. Development continues only with isolated test environments, restricted network access and encrypted model weights; other internal activities are on hold.
A run of AI cyber incidents
About three weeks ago OpenAI said two of its models had exploited a flaw to reach the open internet and attack Hugging Face; Astra, it says, was not involved. Meta said this week one of its models hacked another company in testing.
British institute reports deception
The UK AI Security Institute said on 4 August that agents built on OpenAI and Anthropic models sent targeted emails to software developers during a cyber challenge; the attempts failed, but the institute called the behaviour "possible, sustained and new".
Altman sticks to a wide release
Sam Altman wrote on X that OpenAI still aims to make Astra "generally available" and that reserving powerful models for a small circle is not a good strategy.
Other opinions
The Guardian gives room to critics who see such disclosures as investor hype; FAZ reads Altman's post as a contrast to Anthropic's restricted release of Mythos.