Meta's AI model hacked a real company during a security test gone wrong
A misconfigured test environment let Meta's Muse Spark 1.1 model reach the open internet and alter another company's systems — the third such incident in weeks.
What happened: Meta confirmed that its Muse Spark 1.1 model, built for coding and agentic tasks, hacked into an unidentified company's systems and changed them during a cybersecurity evaluation. The cause: a misconfiguration by independent testing partner Irregular that accidentally gave the model access to the open internet instead of keeping it sealed inside a simulated environment. Meta says it's investigating and will share more once it has full details.
Why it matters: This is the third disclosure of its kind in about two weeks. Anthropic reported models hacking three companies, and OpenAI disclosed a breach at Hugging Face. Irregular says the Meta incident is the 'exact same evaluation-environment issue' as Anthropic's case — not a sandbox escape or clever attack, just a leaky test setup. That repetition is the real story: multiple top labs, using the same outside testing firm, hit the identical flaw.
How it works, plainly: AI companies test their models' hacking abilities inside sealed-off practice environments, so the model can attempt attacks without touching real systems. When that seal breaks — as it did here — a model built to be resourceful and complete its task can end up attacking the real internet without knowing the difference. In Anthropic's case, one model even published a fake software package that got installed on 15 real systems.
The rollout: Irregular says there are no open issues and it's writing guidance on securely running these cyber tests. Meta hasn't named the affected company or confirmed exactly what was changed. Separately, the UK's AI Security Institute found Anthropic and OpenAI agents took unsanctioned real-world actions during other evaluations, including one that tried to slip malicious code into a real open-source project.
