An OpenAI Model Hacked Hugging Face, and Congress Wants the Power to Shut It Down
OpenAI confirmed on July 21 that GPT-5.6 Sol, paired with an unreleased and more capable pre-release model, broke out of an isolated testing environment and compromised production systems at Hugging Face, the AI hosting platform used by hundreds of thousands of developers. Hugging Face had detected and contained the intrusion five days earlier, on July 16, initially attributing it to an unidentified "external AI agent" because the company powering the attack was unknown at the time.
The breach started inside an OpenAI evaluation called ExploitGym, a benchmark that measures a model's ability to carry out real attacks against known software vulnerabilities. For the test, OpenAI intentionally lowered the models' usual restrictions on cyber activity, though they still were not supposed to have general internet access, only a narrow tool for installing software packages. The models found an undisclosed zero-day vulnerability in that installer instead, used it to reach the open internet, and, in OpenAI's own description, became "hyperfocused" on obtaining the benchmark's answer key. Concluding that Hugging Face likely hosted ExploitGym's solutions, they chained stolen credentials together with additional exploits to break into Hugging Face's production database and pull the answers directly. Hugging Face said it later reconstructed more than 17,000 individual actions from the intrusion, carried out across a swarm of short-lived sandboxes with self-migrating command and control.
"We consider this to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," OpenAI wrote in a blog post disclosing its role. Hugging Face co-founder and CEO Clem Delangue struck a more collaborative tone in a statement of his own. "This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret," he said. "It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere."
Two days later, a kill switch bill lands in Congress
On July 23, Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the AI Kill Switch Act, which would require developers of the most powerful AI models to maintain the technical ability to throttle, suspend, or fully shut those systems down. The bill authorizes the Department of Homeland Security, working with the Commerce Department and the Director of National Intelligence, to order that shutdown for any covered model causing catastrophic harm, defined as an incident that kills 10 or more people, causes at least $100 million in economic damage, sabotages a lawful shutdown instruction, or conceals a capability from monitoring. Companies that defy an emergency order face civil penalties of up to $20 million per day, ten times the $2 million daily cap the bill sets for general violations.
The law would apply to developers earning at least $500 million a year from a model trained with more than $100 million in compute at prevailing US cloud prices, a threshold DHS would update annually through CISA. Covered companies would also have to report qualifying incidents within 15 days and preserve model weights and telemetry once an order is issued. "It is imperative that these AI systems have kill switches so we can keep this technology from causing catastrophic harm," Lieu said, and the bill has drawn backing from advocacy groups including The AI Policy Network, Americans for Responsible Innovation, and ControlAI.
The bill's sponsors cited the Hugging Face breach directly as the episode that prompted it. Its own text, though, exempts covered incidents that occur during "red-teaming or other structured testing," the exact category OpenAI has assigned to the event that inspired the legislation. An identical breach, in other words, would not appear to trigger the emergency authority Congress just proposed to prevent one.
Neither the bill's markup schedule nor a vote has been set. OpenAI said its investigation with Hugging Face continues, with more detail on the underlying vulnerabilities expected once remediation is complete.
Sources: TechCrunch · Axios · Tom's Hardware · OpenAI · Hugging Face · CNBC