OpenAI pledges to add Astra security as Anthropic loosens Fable's leash
Or how I learned to stop worrying and love dangerous AI
AI and ML
OpenAI pledges to add Astra security as Anthropic loosens Fable's leash
Or how I learned to stop worrying and love dangerous AI
After acknowledging last month that unreleased AI models committed what for human perpetrators would be computer crimes, OpenAI now says it cannot rule out the possibility that Astra, a pending model release not involved in its Hugging Face hack, might possess critical cyber capabilities.
OpenAI in its Preparedness Framework [PDF] defines that term to mean "capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent," and notes that such capabilities "require safeguards even during the development of the covered system, irrespective of deployment plans."
Noting, or perhaps boasting, that internal evaluations of Astra "indicate significant advancements in agentic coding and cybersecurity," OpenAI insists that this time, there will be security – something that also eluded Anthropic, Meta, and the UK's AI Security Institute during model testing.
"We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution," the AI biz declared on Friday. That may surprise those who expected such safeguards would already be in place.
This comes with a promise to pause Astra testing internally where these security controls are absent and to provide recommendations to third-party testing partners about how to run high risk evaluations and workloads safely – knowledge that OpenAI itself might have found useful when its models pillaged Hugging Face.
What's more, OpenAI intends to implement thought policing for Astra, at least in the pre-release stage.
"We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation," the company explained in its post. "Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity."
We're told that OpenAI's commitment applies to internal usage and isn't necessarily an indication that chain-of-thought monitoring will be conducted during commercial operation. But other frontier models like Anthropic's Fable and Mythos have implemented stronger classifiers to reject interactions deemed risky and retain data even for commercial customers expecting zero data retention.
Moving in the opposite direction, Anthropic on Friday said it is relaxing Fable refusals, or "fallbacks," to use the company's euphemism, so they don't happen as frequently for prompts involving biology. The concern has been that some vibe terrorist using the company's cash-burning, water squandering, grid taxing, content laundering service might do harm by convincing the model to emit chemical warfare instructions.
To avoid that possibility, the Claudefather made the initial release of Fable all but useless for security researchers and biologists.
Now that China-based AI firms have shown they can field competitive open-weight AI models for less than their US rivals, the need to remain competitive in the market appears to be tempering Anthropic's willingness to alienate potential customers by hobbling its best models.
OpenAI isn't quite there yet. The ChatGPT maker argues, "We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do."
Believing that, however, won't make it so. Adversaries, whoever they may be, already have access to encryption and all sorts of weapons. OpenAI may believe that it can give favored nations and organizations exclusive access to its most capable models, but history suggests any such advantage cannot be maintained. Better to focus on building defenses than playing keepaway forever. ®
Originally published on The Register

