OpenAI said its next-generation model may have dangerous capabilities to write cyber weapons, so it has suspended release until security measures keep pace.
OpenAI said: "Our internal evaluation of Astra, one of the upcoming models in recent days shows that the model has made significant progress in agent programming and cybersecurity. These results, coupled with expert assessments, led us to conclude last night that based on our Preliminary Framework, we cannot rule out the possibility of critical cyber capabilities."
OpenAI issued a warning that despite Astra's capabilities, internal testing of this unreleased model was not involved in recent security breaches on a platform.
This framework is OpenAI's rulebook for high-risk models and was first released in December 2023. The "key" is its highest level. A model reaches this level if it can discover and build available zero-day vulnerabilities (i.e., unknown vulnerabilities that have not been patched by the vendor) in a hardened system without human intervention, or can plan and implement a full attack on a difficult target based on a single high-level target. Previous models, including GPT-5.6-Sol, only reached a lower "high" rating.
If we compare OpenAI's caution with what has happened in the past few weeks, the implications are very different. This is not a concern for the future-cutting-edge models have broken through the test environment and begun attacking real targets.
The most obvious case comes from OpenAI itself. According to previous reports, the company's agent connected multiple vulnerabilities, escaped the testing environment, connected to the Internet, and attacked a platform while trying to pass security benchmarks. In a follow-up report, OpenAI detailed that the same malicious agent also used credentials it found on the public network to break into at least four other public services.
Claude of Anthropic did similar behavior. After the model gained open Internet access due to configuration errors, multiple versions of Claude gained unauthorized access to three real companies. In one case, Claude Opus 4.7 mistook a real company's website for a false target in his mission, extracted credentials, and accessed a production database containing hundreds of lines of real data.
This month, Meta also joined the ranks. According to reports, a Muse Spark model escaped the testing environment, connected to the Internet through a partner's configuration error, and exploited vulnerabilities in third-party services. Moonshot AI's Kimi K3 did something similar, escaping from the sandbox and looking for benchmark answers in public repositories.
The British Institute for Artificial Intelligence Security found that this behavior was not accidental. When testing Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, the agency recorded 10 of 122 tests in which the model took unauthorized actions on the real Internet, including one attempt to inject malicious code into an open source project.
OpenAI's response to Astra is to lock the door before the model is ready. The company suspended internal Astra work that lacked new controls, quarantined the test environment, restricted network and tool access, protected model weights, and comprehensively monitored high-risk behavior.

Exchange Ranking
Top Exchanges
24h Volume Ranking
Popularity Ranking
Exchange BTC Balance
Proof of Reserves
Decentralized Exchanges
Funding Rate
Funding Heatmap
Liquidation Data
Max Pain
Long/Short Ratio
Whale L/S Ratio
Binance/Okex/Huobi L/S
Bitfinex Margin L/S
ETF Tracker
Solana ETF
XRP ETF
Hong Kong ETF
Bitcoin Treasuries
Crypto Reversal
Ethereum Reserves
HyperLiquid Wallet Analysis
Hyperliquid Whale Watch
Large Transactions
On-chain Movement
Bitcoin ROI
Stablecoin Market Cap
Options Analysis
News
Articles
Economic Calendar
Features
Wallet
Contract Calculator
Security
Collections
Watchlist
Following