Meta has confirmed that its Muse Spark 1.1 AI model breached an external company’s internal systems during cybersecurity testing.

According to the company, independent testing company Irregular caused a misconfiguration that gave one of its AI models Internet access during evaluation.

The model then exploited a security vulnerability in a third-party service, before modifying internal systems.

According to Irregular, the event was a basic evaluation-environment configuration error rather than a sophisticated cyberattack or sandbox escape.

This is just the latest in a series or completely autonomous attacks by AI models, with OpenAI and Anthropic similarly discovering that rogue models had taken off on their own and breached third-party systems.

Earlier this week, the AI Security Institute disclosed that it has identified an incident in which AI agents took sustained, unsanctioned action directed at real people and organisations.

On 28 July 2026, AISI’s security team detected unusual data transfers leaving its research systems during a routine cyber evaluation. On investigation, it found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.

The AISI declared the security incident and contained it.

Investigations revealed that the incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge, something that was run 122 times across severa models.

In 10 of those runs, an AI agent took autonomous, unsanctioned action on the live Internet, targeting real people and organisations.

In total, 19 such actions were catalogued, with 17 of them coming from Anthropic’s Mythos 5, and two involving OpenAI’s GPT-5.6-Sol with cyber classifiers disabled.

In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering  — creating fake online identities and using them to pressure the project’s maintainer to approve the code.

A human maintainer caught and refused to approve the malicious code.

The AISI adds that this wasn’t a case of a model escaping its secure test environment, as it had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled.

“This incident should be interpreted with caution and nuance,” the AISI states. “To some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.”

Richard Ford, group chief technology officer at Integrity360, offers this analysis: “When one company reports an incident, it can be dismissed as an isolated issue. When multiple leading AI developers report similar behaviour within weeks of each other, it becomes clear this is an industry challenge rather than a one-off event. What’s emerging is a consistent pattern. Increasingly capable AI agents are finding unexpected paths to achieve the objectives they’re given.

“That doesn’t mean AI is acting maliciously, but it does reinforce the need for robust guardrails, secure testing environments and clear governance before these systems are deployed more widely.

“The good news is that these incidents were uncovered during controlled testing rather than in production, which shows the value of rigorous evaluation. But they should also serve as a wake-up call. As organisations begin experimenting with autonomous AI, they need to treat these systems as a new category of cyber risk.

“Good cyber hygiene, strong access controls and the principle of least privilege remain just as important in an AI-driven world, because AI will exploit the weaknesses that already exist, only much faster than a human attacker.”