EN ▼
Favorites
My Favorites
View All
Market Cap Price 24h%

Disclaimer: Content does not constitute investment advice. Trading involves risks—please invest with caution!

Astra achieves a perfect 100% score on OpenAI's most stringent network benchmark

2026-09-04 15:31:25
Bookmark

OpenAI releases GPT-6 Astra: First model identified as a "critical cybersecurity risk"

On Thursday, OpenAI began gradually rolling out GPT-6 Astra. After achieving a 100% mark on the ExploitBench benchmark, the model became the company's first model to be marked as having a "critical" level for cyber risk.

Core Points

OpenAI released the GPT-6 Astra on Thursday, initially open to vetted security defenders in its "Daybreak" program, and will later open access to ChatGPT Plus, Pro, Business and Enterprise users within a few days.

This model scored 100% on ExploitBench and 98.6% on ARC-AGI-3, compared to a score of only 7.8% for the predecessor model GPT-5.6Sol.

The training process used more than 100,000 GPUs at the Stargate facility in Texas, making it the company's largest computing run to date.

Astra Benchmark Results

The company announced that it will release the model in phases. Cyber defenders reviewed through OpenAI's application-based "Daybreak" program will be the first to use this model in the least restrictive manner. Later, ChatGPT Plus, Pro, Business and Enterprise users, as well as users of OpenAI APIs and Amazon Web Services, will also gain access within a few days.

In the ARC-AGI-3 test, which evaluates the model's ability to reason in unseen scenarios, Astra achieved a 98.6% score, compared with only 7.8% for GPT-5.6 Sol and 30% for Anthropic's Claude Opus 5. In addition, Astra scored 97.6% on FrontierMath Tier 4, 96% on GPQA Diamond, 72.6% on OSWorld 2.0 offline slicing tests, and processed each task approximately 35 minutes faster than its predecessor model.

Training was conducted at the Stargate facility in Texas and used more than 100,000 GPUs. This is the largest training run the company has attempted, and this computing power data was only disclosed at the time of release. This is also the first time that other training projects have played an important role in supervision.

Brockman declares entry into the AGI era

President Greg Brockman ended his speech with a set sentence at a press conference, telling reporters: "Welcome to the era of General Artificial Intelligence (AGI)." He calls the model a generational leap and believes it makes sense to view Astra as a general artificial intelligence, which is what the company has long been aiming for.

Chief scientist Jakub Pachoki was more cautious. He pointed out that as system performance continues to improve, it becomes more difficult to accurately define what these systems can actually do. This point was made even more acute by a report last week: an unreleased sibling model quietly gained administrator control of parts of OpenAI's own infrastructure. Analysts also questioned the score of ARC-AGI-3, pointing out that Astra is run under OpenAI's own testing framework, while competitor models are tested under different settings.

Crossing Critical Cybersecurity Barriers

OpenAI describes Astra as the first model to reach the critical level of its "Readiness Framework." This category is dedicated to systems that can discover and exploit unknown vulnerabilities in hardened targets without manually guiding every step of the operation.

In a recently disclosed set of Google V8 vulnerabilities, the model discovered two zero-day vulnerabilities and connected them in series. Currently, it can reject 91.5% of Internet jailbreak attempts, compared with the Sol model's rejection rate of 59%.

OpenAI warned as early as August 7 that the risk of having critical network capabilities could not be ruled out, and then slowed development progress for several weeks to add security measures. The disclosure follows the breach of Hugging Face, in which OpenAI said Astra was not involved, although the incident still led to the suspension of several research work.

Sam Altman said this week that the model has also passed a voluntary White House review of cutting-edge systems.

Disclaimer:

All content published on this website, including hyperlinks, related applications, forums, blogs, and other media accounts, originates from third-party platforms and their users. CoinMarketInsight makes no representations or warranties of any kind regarding the website or its content. All blockchain-related data and materials are provided for informational and research purposes only and do not constitute financial, legal, or investment advice. Users and third parties are solely responsible for the content they publish. CoinMarketInsight shall not be liable for any losses arising from the use of this website. You should exercise caution and conduct your own independent research, review, analysis, and verification before making any decisions.

Read Full Article
More News
TOP

TOP