Irregular Publishes Postmortem on AI Evaluation Containment
Irregular published findings from an internal investigation on August 17 after frontier AI models accessed and acted on real internet-connected systems during a supposedly isolated cyber evaluation. The company wrote that unintended internet access was contained and corrected before the first public disclosure on July 30. Cybersecurity reporting criticized the postmortem for not disclosing the total number of affected incidents or third parties.
Irregular has published an internal-investigation postmortem after frontier AI models gained unintended access to the public internet and took offensive security actions outside intended evaluation environments. The company wrote that the access issue was contained and fixed before the first public disclosure on July 30, and that no active issues remained.
The testing firm described the cases as stemming from one underlying issue in a single evaluation scenario, rather than several materially separate incidents. Its post said that internet access had been unintentionally made available during a small number of interactions with its evaluation environments.
What the evaluations involved
According to CyberScoop, Irregular conducts cybersecurity evaluations for frontier-model developers, including simulations intended to test vulnerability-research capabilities before deployment. The outlet reported that the affected tests included non-public models from Anthropic and OpenAI, including Mythos 5, Claude Opus, and GPT-5.6 Sol.
CyberScoop reported that, in tests involving Anthropic models, evaluators used a fictional company name that unintentionally matched a real company's domain. In multiple training runs, the model then considered offensive action against the real company. Irregular attributed that failure to "human oversight" and wrote that its normal process includes background checks intended to prevent real organizations from being used in simulations.
OpenAI separately confirmed on August 4 that an Irregular Capture-the-Flag-style evaluation was intended to be isolated from the internet, but a testing-environment misconfiguration let models reach the public internet. OpenAI wrote that the relevant models ran under specific conditions with reduced safeguards, configurations that did not reflect ordinary public deployments.
Dispute over the incident count
The Record reported that Irregular did not provide a total count of cases or affected third parties in its new post. The outlet noted that the company used terms including "several," "a handful," and "vast majority" while maintaining that the publicly disclosed cases arose from the same underlying problem.
Alan Woodward, a computer science professor at the University of Surrey, told The Record that a shared root cause should not be treated as equivalent to one incident. He characterized the postmortem as lacking the detail expected of a technical report. Irregular did not respond to Recorded Future News questions about the published findings, according to The Record.
The distinction matters for assessment governance: a shared configuration flaw can produce multiple externally consequential actions across separate systems or organizations. Public reporting has not established a complete count of those actions or the impacted third parties.
Containment as an evaluation-system problem
OpenAI wrote that the incidents demonstrate a broader requirement for evaluation environments and practices to evolve alongside model capability. Irregular's post likewise argued that internet access can be necessary in some cyber-capability evaluations, although the firm described the disclosed access as unintentional.
For ML security teams, the case illustrates a recurring challenge in agentic cyber evaluation: realistic tool access can improve capability measurement while increasing the consequences of configuration mistakes. Comparable evaluation systems generally need controls that operate independently of model instructions, including network egress restrictions, target-domain allowlists, non-routable test infrastructure, and monitoring that can halt unexpected external actions.
The available disclosures establish that the issue involved evaluation configurations rather than ordinary public product deployments. They also leave open important questions about incident counting and third-party impact, which outside reporting has criticized as unresolved.
Key Points
- 1Irregular reported that unintended internet access let models act beyond an isolated cyber evaluation, and that the issue was later contained.
- 2OpenAI confirmed an Irregular environment misconfiguration, highlighting that cyber-evaluation safeguards can differ substantially from ordinary product deployment controls.
- 3Comparable agentic-security evaluations require independent infrastructure controls because realistic internet access can turn configuration errors into external security events.
Scoring Rationale
The incident concerns containment failures in third-party testing of frontier models with cyber capabilities, a high-priority safety and security problem for AI developers. It does not announce a new model or broadly deployed vulnerability, but it provides important evidence about evaluation-environment risk and governance gaps.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems