Meta has disclosed that one of its artificial intelligence (AI) models gained access to another organisation’s systems during a security evaluation, becoming the latest technology company to report an AI hacking incident during testing.
The Facebook parent company said the incident occurred during an independent assessment carried out by AI security firm Irregular. According to Meta, the model’s access resulted from a misconfiguration in the testing environment rather than a flaw in the AI system itself.
According to a recent report by the BBC, Irregular, which also conducted security tests on Anthropic’s AI models, said the Meta incident stemmed from “the exact same evaluation-environment issue” that had already been identified during Anthropic’s testing. The company said it is preparing guidance on how AI agents should be securely evaluated in cyber-security exercises.
The disclosure follows a series of reports from OpenAI and Anthropic over the past two weeks. OpenAI said some of its AI agents attempted to attack publicly available online services, including the AI platform Hugging Face, during internal evaluations.
Anthropic subsequently carried out similar tests and found that one of its Claude models had also accessed systems belonging to several organisations after a testing misconfiguration granted it internet access.