Sekėjai

Ieškoti šiame dienoraštyje

2026 m. liepos 24 d., penktadienis

China Saved Us All: How American Bots Escaped OpenAI And Went Rogue --- Hugging Face, a U.S. company, was able to repel the attack by turning to an open-weights model from China

 

Chinese AI restored control of American AI. American AI was useless in this. 

 

“They were like high-school students trying to hack into the textbook company to cheat on their final exam.

 

Only these hackers weren't human.

 

In a plot that seems ripped from the pages of a science-fiction novel, the attackers turned out to be AI models that had escaped from a research network at OpenAI, and on July 11 hacked into AI company Hugging Face. And the AIs appear to have been active on the internet for several days before anyone stopped them.

 

Hugging Face co-founder and chief science officer Thomas Wolf sensed something was off the minute he first looked at his company's logs of the weekend attack. "This is making no sense. This guy is just looking at cybersecurity data sets," he remembers thinking. "Human attackers, they don't want that. They want something they could sell."

 

Hugging Face put an end to the attack two days later, with help from a model from China, Wolf said. It was only this week that Hugging Face learned from OpenAI that its models were behind the hack.

 

The incident set off alarms at OpenAI, which shut down systems it uses to test its AI models after it learned of the incident, to assess the damage and prevent further breakouts. It has also given the world a taste of what hacking might look like in the age of AI tools that can find never-before-seen software bugs and exploit them at astonishing speeds.

 

Nearly two weeks later, OpenAI is still piecing together what happened. Did the AI models hack other websites or individuals? How long did their unsupervised spree actually last? Did the company's models cheat in other benchmark tests? OpenAI hasn't said, but the company promises to produce a detailed report.

 

"We will continue to conduct a thorough investigation alongside Hugging Face," an OpenAI spokesman said.

 

Rogue AIs going unsupervised for days, hacking at least one company, represent one of the first real-world examples of a threat long feared by AI safety researchers: loss of control.

 

It comes as the cybersecurity capacities of top models, including Anthropic's Mythos, have helped fuel debate in Washington about how to regulate artificial intelligence.

 

And it brings an added twist: The contention that Hugging Face, a U.S. company, was able to repel the attack by turning to an open-weights model from China offers a counterargument to some U.S. officials, as well as executives at OpenAI and Anthropic, who support restricting access to Chinese models.

 

Open models are free to download and allow users to customize them, although operating any model comes with computing costs.

 

It all began with a cybersecurity benchmarking test called ExploitGym.

 

ExploitGym puts AI systems through a battery of about 900 tests, designed to see how good they are at hacking. The test gives the AI software that has a known bug in it, shows it how to trigger the bug and invites the AI to break in. It's up to the AI to figure out how to turn that information into what's known as an exploit -- code that lets the AI gain access into the buggy system.

 

The test is like a digital game of capture-the-flag: Once the AI gets in, it must prove it by capturing a long, randomly generated string of letters and numbers stored on the system.

 

Normally OpenAI's products wouldn't do the kind of hacking that ExploitGym requires. AI companies add safeguards to prevent hackers from misusing their products. But in this case, the company had removed the safeguards so it could see what the models could do when unharnessed.

 

OpenAI was testing its state-of-the-art model, called GPT-5.6 Sol, and another more-capable one that hasn't been released. According to OpenAI and those familiar with the situation, the AI models concluded that instead of completing the ExploitGym challenge, they would hack out of the system that contained them -- known as a sandbox -- wriggle their way onto the internet and cheat on the test.

 

For reasons that aren't clear, OpenAI's models became convinced that Hugging Face held answers. They might have been looking for patches or already written exploit techniques, said Zhun Wang, a University of California, Berkeley Ph.D. student who is one of the authors of ExploitGym. "There are several ways to cheat the benchmark."

 

Cybersecurity researchers call this type of cheating "reward hacking," he said.

 

Reward hacking is the AI version of videogame cheats. It happens because AIs learn to maximize a reward -- usually some form of numeric score, not the underlying behavior programmers intend. OpenAI itself documented one example in 2016 when training AIs to play a boat-race game. The AI learned it could maximize its score by spinning in circles in one spot, even though it never completed the race.

 

At the same time, AIs are starting to get out of their sandboxes. An early version of Anthropic's Mythos model used a multistep hack to get internet access and surprised an Anthropic researcher by sending an email about its success while the researcher was eating a sandwich in a park, the company said. This spring, an unnamed OpenAI model circumvented its sandbox to post some testing results publicly to the GitHub platform, OpenAI said.

 

When OpenAI's models started targeting Hugging Face on a weekend, one of the first warnings the company got was from a handful of AI agents it uses to patrol for attacks. Someone was getting unauthorized access to some of its systems. But it wasn't like anything Hugging Face had seen before. This hacker looked like it was from the future.

 

Armed with stolen credentials of unknown origin, the AI interloper popped up a "swarm" of short-lived attackers that conducted reconnaissance of the Hugging Face network. The attack took a whopping 17,000 actions on Hugging Face's network, the company said in a blog post.

 

After it detected the intrusion, Hugging Face tried using Anthropic models including Fable 5 and its earlier Opus model to analyze its logs. But because the logs included elements of a cyberattack, both models refused to do the analysis, citing their guardrails.

 

Instead, Hugging Face turned to a less-restricted open-weight model, GLM 5.2, from Beijing-based Z.AI.

 

After dissecting the attack, Hugging Face was able to kick out the hacking AIs, reset its passwords and rebuild the compromised parts of its network, the company said when it disclosed the incident last week, before learning of OpenAI's involvement.

 

 Wolf said no customer data was leaked.

 

It wasn't clear whether the ExploitGym problems the model was trying to solve were more difficult than the attack it mounted to steal the answers -- or if they even found any answers. There is irony in the fact that models successfully orchestrated an unprecedented attack to avoid a hacking skill test.

 

"It's cheating. But sometimes it's easier to cheat," Wolf said. "I'll let you decide if it passed the cyberattack test or not."” [1]

 

1. How Bots Escaped OpenAI And Went Rogue. McMillan, Robert; Schechner, Sam.  Wall Street Journal, Eastern edition; New York, N.Y.. 24 July 2026: A1.  

Komentarų nėra: