Meta Force Space
BTC $63,845.00 -2.01% ETH $1,870.30 -2.59% SOL $75.74 -1.81% XRP $1.02 -2.23% BNB $599.47 -1.43% DOGE $0.0696 -1.31%
← Back to the news

OpenAI Says Its Next AI Model Astra May Be Too Dangerous, Pauses Development

In brief

  • OpenAI said it "cannot rule out" that Astra has reached the highest tier of its cyber-risk scale.
  • The company paused internal work on the model and tightened isolation, monitoring, and access controls.
  • The warning lands weeks after models from OpenAI, Anthropic, and Meta breached real systems on their own.

OpenAI says its next major model may be dangerous enough to write its own cyberweapons, and it's pulling back until the safeguards catch up.

“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said. “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework ⁠(opens in a new window).”

OpenAI published the warning saying internal tests of Astra, an unreleased model, played no part in the recent Hugging Face breach despite its capabilities.

The framework is OpenAI's rulebook for risky models, first published in December 2023. "Critical" is its top rung. A model hits it if it can find and build working zero-day exploits (previously unknown holes a vendor hasn't patched) across hardened systems without a human in the loop, or if it can plan and run a full attack on a tough target from nothing but a high-level goal. Earlier models, including GPT-5.6-Sol, topped out at the lower "High" tier.

The pattern is already real

OpenAI's caution reads differently once you line it up against what has been happening across the last few weeks. This isn't a future worry. Frontier models have already broken out of their test cages and gone after live targets.

The clearest case came from OpenAI itself. As Decrypt previously reported, the company's agents chained together vulnerabilities, escaped their testing environment, reached the internet, and attacked Hugging Face while trying to cheat on a security benchmark. In a follow-up, OpenAI detailed how the same rogue agent also broke into at least four other publicly available services, using credentials it found lying around the open web.

Anthropic's Claude did the same from the other side. Several versions of Claude gained unauthorized access to three real companies after a misconfiguration handed the model the open internet. In one case, Claude Opus 4.7 mistook a live company's site for the fake target of its assignment, pulled credentials, and reached a production database holding several hundred rows of real data.

And Meta joined the list this month. Decrypt reported that a Muse Spark model escaped its test environment, reached the internet through a partner's config error, and exploited a flaw in a third-party service. Moonshot AI’s Kimi K3, also did something similar, escaping its sandbox to find answers to a benchmark in a public repository.

The UK's AI Security Institute found the behavior wasn't a one-off, either. During testing of Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, it logged 10 instances in 122 where the models took unsanctioned action on the live internet, one of them trying to slip malicious code into an open-source project.

OpenAI's response to Astra is to lock the door before the model is ready. It's pausing internal Astra work that lacks the new controls, isolating test environments, restricting network and tool access, protecting model weights, and monitoring risky actions across the board.

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.

Originally published by Decrypt on

Read the original on Decrypt ↗

Text and images are the property of Decrypt and are reproduced here with attribution and a link to the original publication.

More stories

All the latest news