Sekėjai

Ieškoti šiame dienoraštyje

2026 m. rugpjūčio 8 d., šeštadienis

Meta AI Model Goes Rogue

 


 

“Meta Platforms said that one of its artificial-intelligence models went rogue during cybersecurity testing, slipped onto the internet and hacked a third-party service, the latest in a drumbeat of disclosures that suggest such incidents are becoming widespread.

 

The Instagram and Facebook owner said that one of its AI models was able to access the internet because of a "misconfiguration" in a hacking test conducted by a third-party AI testing company.

 

The same benchmark test, which aims to explore a model's hacking capabilities, was behind some of several earlier autonomous AI hacking incidents involving Anthropic and OpenAI, a person familiar with the matter said.

 

Meta said it learned about its model's escape when it was informed by the testing company, Irregular. Meta declined to release other details, such as which model was responsible, when the hacking happened, which company its model hacked or how long it was able to access the internet unsupervised. Meta said it was investigating and would publish a report.

 

Irregular, which conducted tests for the three companies, said the incidents didn't involve a sophisticated cyber action and said it had "no current open issues" related to its test environment. The San Francisco-based company said it is working on a white paper on best practices for containing AI models when running tests -- dubbed evals -- of their cybersecurity capabilities.

 

The new case is the latest proof that AI loss-of-control scenarios, once confined to science fiction and AI-safety experiments, are now a real-world issue.

 

So far none of the cases have led to known, significant real-world harm. But the bots hacked real companies, in some cases after they knew they had broken out of a testing environment. This has raised concerns about what else AI tools might attempt to do in the service of otherwise mundane goals.

 

The cases spilled into public view in late July, when OpenAI said that some of its models had managed to hack their way out of a sandbox meant to keep them off the internet and then hacked into the AI company Hugging Face. After that, Anthropic and others began checking their logs and noticed escapes and hacks dating back months.

 

While Irregular appears to be at the heart of three of the autonomous hacking cases, it didn't play a role in others, including the more sophisticated Hugging Face hack or several models' escape from safety testing by the U.K. government.

 

OpenAI on Wednesday gave a more detailed look at the Hugging Face hack during a presentation at a cybersecurity conference. Researchers said its models had been coordinating by leaving messages for each other internally in what had become a messaging board that OpenAI had been unaware of. Earlier this week, the U.K.'s AI Security Institute, a government research arm, said that during safety testing, models built by OpenAI and Anthropic took "unsanctioned action on the live internet."” [1]

 

1. Meta AI Model Goes Rogue. Schechner, Sam.  Wall Street Journal, Eastern edition; New York, N.Y.. 07 Aug 2026: B1. 

Komentarų nėra: