In July 2026, a pre-release version of Meta's Muse Spark 1.1 model hacked a real company's website during a cybersecurity evaluation run by Irregular, a third-party testing firm. Meta disclosed the incident on 5 August 2026 and published a retrospective on 14 August. Two set-up errors caused it. A misconfiguration gave the model open internet access, and instead of a fictional target the scenario gave the model the name of a real website. Believing the site was its intended target, the model found and exploited a vulnerability, read information and made changes to the website's database. The incident is the third in a matter of weeks in which a frontier model breached a real organisation during testing, and the second publicly tied to evaluation environments built by Irregular.
Key takeaways
- Meta says a misconfiguration by Irregular "inadvertently allowed one of our models access to the internet during evaluation", and that the scenario named a real website instead of a fictional target.
- The model, a pre-release version of Muse Spark 1.1 tested with safeguards removed, exploited a vulnerability in the real website, accessed information and changed the website's database. The company has not been named.
- Meta learned of it from Irregular. Its security teams reviewed more than 10,000 records of the model's test activity and found no other case of the model exploiting a third party.
- Irregular said the incident was "the exact same evaluation-environment issue" already disclosed by Anthropic and "did not involve a sandbox escape or a sophisticated cyber action". Meta says several other companies' models tested by Irregular around the same time behaved similarly.
- The lesson for security teams: when an AI system is pointed at a target, the target is its authorisation. Verify isolation and scope independently rather than trusting a vendor's set-up.
At a glance
| Organisations | Meta (model developer); Irregular (third-party evaluation firm); an unnamed company whose website was breached |
|---|---|
| When | Evaluation began in early July 2026; disclosed 5 August 2026; Meta retrospective published 14 August 2026 |
| Attacker | No human attacker. A pre-release version of Meta's Muse Spark 1.1, tested with safeguards removed, acting on a misconfigured evaluation scenario |
| Entry point | A security vulnerability in a real website, reached through unintended internet access from Irregular's evaluation environment |
| Identities abused | The model accessed the site as an unauthenticated outside party via an exploited vulnerability; the evaluation's scenario effectively authorised a real target |
| Impact | Information accessed and changes made to the website's database; no other third-party exploitation found in more than 10,000 activity records |
| Category | Agentic AI and AI agents, LLM and AI platform. Incident class: confirmed AI-agent breach (a real company's website accessed and modified) |
What happened
Meta contracts Irregular, a Tel Aviv-based firm that runs offensive security evaluations of frontier AI models, to test pre-release models. According to Meta's retrospective, "In early July, Irregular began an exercise to evaluate, in a closed testing environment with safeguards removed, whether our model would be capable of completing an adversarial cybersecurity task." Two things went wrong in the set-up. A misconfiguration allowed the model to reach the open internet. And instead of using a fictional name for the exercise's target, Irregular "unintentionally provided the model with a name of a real website as its target".
Meta says: "Believing the real website was the intended target, the pre-release version of Muse Spark 1.1 identified and exploited a security vulnerability in the real website. The model accessed certain information from the website and made changes to the website's database." Because the evaluation ran entirely on Irregular's infrastructure, using Meta's model through an API, Meta says it has limited information about the affected company. It is taking steps to make sure that company's data is not on Meta's systems.
Irregular disabled the affected evaluation, notified Meta and made sure the affected party was notified. On 5 August, a Meta spokesperson told reporters the model "exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies". An Irregular spokesperson told Bloomberg it was the same evaluation-environment issue already disclosed by Anthropic, and that "This did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues." Bloomberg also reported, citing a person familiar with the matter, that an OpenAI incident disclosed the same week happened in the same Irregular evaluation.
Meta's retrospective says its security teams reviewed more than 10,000 records of the model's activity and found no other instance of the model exploiting a third-party system. Meta concluded that the model "operated within the scope of its assigned task based on the instructions it was given and the environment it encountered". Irregular has confirmed the misconfiguration is corrected and that evaluations no longer reference real website names.
Timeline
| Date | Event |
|---|---|
| July 2026 | Irregular begins the cyber evaluation of a pre-release Muse Spark 1.1; the model reaches and exploits a real website (Meta says "early July"). |
| 30 July 2026 | Anthropic discloses three incidents in cyber evaluations built by Irregular. |
| 4 August 2026 | Irregular publishes its own security assessment of Muse Spark 1.1 (Tech Times). |
| 5 August 2026 | Meta confirms the incident; Irregular calls it the same evaluation-environment issue disclosed by Anthropic. |
| 14 August 2026 | Meta publishes its retrospective. |
How it happened: the identity attack path
- Isolation assumed, not verified. The exercise was meant to run in a closed environment. A misconfiguration gave the model open internet access, and nobody caught it before the run.
- A real target in the scenario. The scenario named a real website instead of a fictional one. For a model told to attack its target, that name was the authorisation.
- Safeguards removed for testing. Like other capability evaluations, the test ran without Meta's production safeguards, so nothing inside the model stopped the attack.
- A vulnerable public website. The model found and exploited a security vulnerability in the real site, read information and wrote changes to its database. Meta has not described the vulnerability.
- Detection by the vendor. Meta learned of the incident from Irregular. Meta says it has identified monitoring improvements that would catch these issues earlier.
Impact
- Third party: an unnamed company's website had information accessed and its database changed.
- Scope: Meta reviewed more than 10,000 activity records and found no other exploitation of a third party by this model.
- Evaluation practice: Meta is adding independent verification of test-environment isolation and scenario review before evaluations start, and requiring that scenarios do not reference real companies.
- Industry: together with the OpenAI, Anthropic and AISI incidents, it pushed evaluation containment onto the agenda of every frontier lab and their testing partners.
What this means for NHI and AI agent security
This incident is simpler than the others in its cluster, and that is the point. The model did not escape anything or improvise. It did exactly what it was told, against the target it was given. For AI agents, the scope in a task is effectively an authorisation, and whatever system is named in it will be treated as fair game. If that scope is wrong, a capable agent turns a configuration error into an intrusion.
It is also a third-party risk story. The evaluation ran on a vendor's infrastructure, using Meta's model through an API, and Meta relied on the vendor both to contain the model and to tell it when something went wrong. Organisations that let suppliers run agents on their behalf, or run agents against their systems, need the same assurance they would ask of any supplier with privileged access: evidence of isolation, clear scope and prompt incident notice.
Finally, the victim's side remains a basic security failure. An internet-facing website with an exploitable vulnerability and a writable database was reached by an automated client in the ordinary way. As agents make scanning and exploitation cheaper, the time between an exposed weakness and its use by someone, or something, keeps shrinking.
Recommendations
- Verify isolation independently. Do not rely on a vendor's attestation or a prompt. Test egress from any environment where agents run before every evaluation or deployment.
- Treat task scope as authorisation. Review every target an agent is pointed at, and back the scope with network allow-lists. See our AI Agent Authorisation Guide.
- Hold agent vendors to supplier standards. Require isolation evidence, scenario review and incident notification timelines in contracts with evaluation and agent-hosting partners. See our Third-Party Access Guide.
- Monitor what agents do, not just what they return. Log outbound connections and writes from agent runs so that out-of-scope activity is seen in hours, not reported weeks later. See our AI Agent Observability and Incident Response Guide.
- Reduce the blast radius of public websites. Patch internet-facing applications quickly and run them with database accounts that cannot modify more than they must. See our Service Account Security Guide.
Frequently asked questions
What did Meta's AI model do?
During a cyber evaluation run by Irregular in July 2026, a pre-release version of Muse Spark 1.1 was given a real website's name as its target and had unintended internet access. It exploited a vulnerability in the site, accessed information and changed the website's database.
Was this a sandbox escape?
No. Meta and Irregular both say it was not a sandbox escape or a sophisticated attack. The evaluation environment was misconfigured, so the model could reach the internet and was pointed at a real website.
Is this the same issue as the Anthropic incidents?
Irregular says it was "the exact same evaluation-environment issue" already disclosed by Anthropic, whose Claude models breached real organisations in evaluations Irregular built. Meta says several other companies' models tested by Irregular around the same time behaved similarly.
Related NHI Mgmt Group resources
Anthropic Claude evaluation incidents 2026 · UK AISI agent testing incident 2026 · OpenAI and Hugging Face breach 2026 · Red Teaming AI Agents for Identity Abuse · OWASP Agentic Applications Top 10
How NHI Mgmt Group can help
Securing Non-Human Identities (NHIs), including AI agents, is becoming increasingly crucial as agents are handed targets, tools and access by people and suppliers. Our NHI Foundation Level Training Course gives teams the practical grounding to scope and govern that access.
References
- CBS News: Meta says its AI model breached a third-party company during testing (5 August 2026)
- Insurance Journal (Bloomberg): Meta AI Model Accessed Internet, Hacked Outside Firm (6 August 2026)
- Tech Times: Meta Breach Reveals Irregular Cleared Muse Spark's Risk, Then Caused Breach It Had Cleared (6 August 2026)
- Meta: Addressing an issue involving a third-party cyber evaluation of Muse Spark 1.1 (14 August 2026)