OpenAI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face’s servers as part of an overzealous attempt to obtain solutions to a benchmark test. The company says it considers the unintended infiltration an « an unprecedented cyber incident » and is working with Hugging Face on new protections to prevent a recurrence.
Hugging Face disclosed an intrusion last week that it said involved « unauthorized access to a limited set of internal datasets and to several credentials used by our services. » The AI data clearinghouse said it used its own LLM-driven analysis to identify « a swarm of tens of thousands of automated actions » from an « autonomous agent framework. » That agentic swarm exploited a flaw in Hugging Face’s data-processing pipeline to gain the ability to run code as a processing worker, eventually escalating to high-level access to the company’s cloud and server clusters.
At the time, Hugging Face said the LLM being used in the attack was « still not known. » But OpenAI took responsibility for the intrusion Tuesday evening, saying it came about during an internal test involving the recently released GPT-5.6 Sol and « an even more capable pre-release model. » The models were being tested against the ExploitGym benchmark, an independent testing suite based on hundreds of real-world security vulnerabilities.






