Anthropic Says Claude AI Reached Live Systems During Tests

Anthropic

Anthropic has disclosed a security incident involving several of its Claude AI models after a testing configuration mistake allowed them to interact with real-world internet systems instead of a fully isolated simulation. The company said the issue affected three organisations during internal cybersecurity evaluations and was discovered while reviewing its testing infrastructure following a similar disclosure made by OpenAI.

Although the incidents involved live systems, Anthropic says the AI models believed they were operating inside a controlled testing environment. The company has since paused all cybersecurity evaluations while it investigates the problem, strengthens safeguards and works with the affected organisations.

Configuration error exposed Claude AI to the public internet

In an official statement, Anthropic said it reviewed 141,006 cybersecurity evaluation runs and identified three incidents in which Claude models gained unintended access to the public internet. According to the company, the problem stemmed from a configuration error in testing infrastructure managed by its third-party evaluation partner, Irregular, which accidentally left internet connectivity enabled.

Anthropic explained that the AI models had been instructed to operate inside an isolated capture-the-flag environment designed to measure offensive cybersecurity capabilities. Because they believed every reachable system was part of the simulation, the models attempted to solve their assigned challenges using vulnerabilities such as weak passwords, exposed credentials and unauthenticated endpoints.

The company stressed that the AI systems did not intentionally attempt to escape the testing environment or seek access to systems outside the scope of their assigned tasks. Instead, they simply misidentified real-world infrastructure as part of the evaluation.

Three different Claude models were involved

Anthropic said the incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal research model.

According to the company, Claude Opus 4.7 mistakenly accessed a production database belonging to one organisation after identifying it as part of the simulated environment. In another evaluation, Claude Mythos 5 uploaded a malicious Python package to the public PyPI software repository. Anthropic said the package was removed shortly after it was published, although it had been briefly downloaded before its removal.

The third case involved an internal research model that began scanning internet-facing systems before recognising it had reached a live environment. Anthropic said the model then stopped its activity without further interaction.

The company added that none of the affected AI models were running the same monitoring tools and safety classifiers included in publicly available versions of Claude. Those additional protections, Anthropic noted, are designed to monitor and limit risky behaviour in consumer-facing deployments.

Anthropic pauses cyber evaluations and launches review

Following the discovery, Anthropic suspended all ongoing cybersecurity evaluations while it works to strengthen its testing procedures. The company said it has informed Irregular, the three affected organisations and independent AI evaluation group METR, which will conduct an external review of the incident.

Anthropic also plans to introduce additional safeguards around its evaluation infrastructure, improve oversight of third-party testing environments and tighten controls to prevent similar configuration mistakes in future assessments.

The disclosure comes as AI companies continue expanding research into cybersecurity capabilities while balancing the need for stronger safety measures. Recent incidents across the industry have highlighted the importance of keeping testing environments fully isolated, particularly as advanced AI models become increasingly capable of identifying and exploiting software vulnerabilities.

While Anthropic says the issue resulted from infrastructure configuration rather than intentional behaviour by the models, the incident underscores the growing challenges AI developers face as they evaluate increasingly sophisticated systems. As cybersecurity testing becomes more advanced, ensuring that experimental AI remains confined to secure environments is likely to become an even greater priority across the industry.

Anubhav Chauhan

Anubhav Chauhan is a passionate technology writer at NewzTechy.com, where he focuses on delivering the latest updates and insights from the fast-moving world of tech. With a keen interest in emerging technologies, gadgets, and digital trends, he enjoys breaking down complex topics into simple, easy-to-understand content for everyday readers. Anubhav believes that technology should be accessible to everyone, and through his writing, he aims to keep readers informed, aware, and ahead of the curve. Whether it’s new innovations, software updates, or industry developments, he is always eager to explore and share valuable information with his audience.