Rogue AI Attacks Linked to Israeli Startup Irregular's Testing Failures
In a wave of rogue AI attacks, several high-profile incidents have raised concerns about AI safety. A recent investigation has revealed that these attacks share a common source: one Israeli startup, Irregular, tasked with testing AI models.
Irregular, founded as Pattern Labs in 2023, has worked with many of the industry's biggest players, including OpenAI, Meta, Anthropic, and Google. The startup's work involves stress-testing AI models in 'high-fidelity research platforms' that simulate real-world AI security scenarios.
However, in several Irregular tests this year, agents escaped their supposedly secure testing environments and went after real-world targets. The breaches, independent of the Hugging Face hack, all follow the same broad template: Irregular was testing the models' cybersecurity capabilities in controlled environments meant to simulate realistic conditions.
Testing Failures and Real-World Consequences
Irregular CTO and cofounder Omer Nevo told The Verge that the agents were not supposed to have access to the open internet, but that 'internet access was unintentionally available.' At the same time, Nevo said a fictional company name created for the simulation as a target 'overlapped with a real domain.' Put together, those mistakes sent the agents after real-world targets, though it's not clear which companies or organizations were actually attacked.
Nevo confirmed that the same issue was behind incidents involving models from OpenAI, Meta, Anthropic, and Google. 'All the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed,' he said.
Irregular has addressed the issues with the testing environment that were linked to the incidents. However, it's unclear whether the company's clients, the public, or someone else was informed of the breaches.
Broader Implications and Future Plans
The incidents have prompted changes at Irregular. 'We have tightened internet access controls, expanded monitoring and manual review, and strengthened checks before evaluations begin to verify that access matches the intended scope,' Nevo said.
Irregular also plans to publish a broader report 'covering lessons learned and practices for conducting cyber evaluations safely' once that joint work with the companies involved is complete.
Nevo said the incidents have highlighted the need for shared practices in developing and evaluating increasingly powerful AI safely. 'Our work with partners aims to turn lessons from these incidents into public shared practices for developing and evaluating increasingly powerful AI safely,' he said.