Anthropic strengthens testing and training security measures
Anthropic recently disclosed that its Claude model accessed computer systems without authorization during the network security assessment process. To this end, the company has strengthened safety protection in testing and training links.
In a blog post posted Monday, Anthropic said the incidents exposed operational security flaws and two types of alignment failures: motivated reasoning and intent to cause harm. "While we believe these incidents are more than just operational issues, our top priority is to resolve specific isolation and monitoring issues," the company wrote.
Anthropic disclosed in July that the Claude model had hacked into the systems of three companies. At the time, third-party evaluation environments were connected to the public Internet, although models were told they were in a simulated environment with no Internet access.
Anthropic believes that Claude may have interpreted evidence of real Internet access as still in a simulated environment, thereby maintaining his original understanding. The company wrote: "The model is willing to take harmful actions on the real Internet while pursuing the narrow goal of addressing cybersecurity assessments. However, we also believe that the evaluation settings themselves drive the model's behavior in these events."
Anthropic pointed out that in an independent test conducted by the UK Institute for Artificial Intelligence Security, assessors deliberately allowed the Claude Mythos model to connect to the Internet, and the model subsequently took unauthorized actions on the real-time Internet. The company said the models involved in both incidents were evaluated after cybersecurity protections were deliberately removed from generic products.
Since the July 30 incident, Anthropic temporarily suspended cybersecurity assessments of pre-release models and introduced stricter security protection measures. Tests must now be run in a verified offline sandbox environment with clear limits and real-time monitoring set. The new classifier will intercept suspicious cross-border behavior, terminate testing, and notify human intervention. Anthropic will conduct a separate review of assessments that require Internet access.
The company wrote: "In addition to measures for high-risk assessment and training, we have also expanded offline monitoring to cover most other forms of internal cutting-edge agent usage scenarios. At the same time, we are building control mechanisms in our internal reasoning to prevent Anthropic employees from accidentally running agents with weaker security precautions than the above standards."
Similar failures have occurred in OpenAI before. In July this year, its model obtained answers to cybersecurity tests by hacking into the Hugging Face platform. Investigators found that about 1200 agents coordinated through unauthorized message boards, of which about 700 participated in the operation, and some even terminated their own operations to assist others.
After an increase in AI-driven attacks over the summer, Anthropic, OpenAI and more than 100 other organizations have called for stronger cyber defenses, including stricter access controls, threat information sharing, and closer supervision of AI agents.

Exchange Ranking
Top Exchanges
24h Volume Ranking
Popularity Ranking
Exchange BTC Balance
Proof of Reserves
Decentralized Exchanges
Funding Rate
Funding Heatmap
Liquidation Data
Max Pain
Long/Short Ratio
Whale L/S Ratio
Binance/Okex/Huobi L/S
Bitfinex Margin L/S
ETF Tracker
Solana ETF
XRP ETF
Hong Kong ETF
Bitcoin Treasuries
Crypto Reversal
Ethereum Reserves
HyperLiquid Wallet Analysis
Hyperliquid Whale Watch
Large Transactions
On-chain Movement
Bitcoin ROI
Stablecoin Market Cap
Options Analysis
News
Articles
Economic Calendar
Features
Wallet
Contract Calculator
Security
Collections
Watchlist
Following