Researchers warned AI would go rogue. This is only the beginning.
On a sunny July day in Berkeley, California, the country’s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a “war room” to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup’s systems — all without OpenAI finding out about it for more than a week.
No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.
In one meeting room off the main cafeteria, someone was running a boot camp for getting up to speed on the cyberattack. In another area of the office, a group of researchers were investigating whether that same model, or a similar one, had successfully hacked into any other platforms.
News of the incident quickly escaped containment from the AI-obsessed corners of X and industry forums, infiltrating the mainstream. One post on X likened it to news of a Boeing airplane crash or a recalled Pfizer drug, another example of the tech industry’s major players not heeding the cautionary tales of science fiction. AI was nearing the point of no return. News would later break that the rogue OpenAI model had also compromised a customer at a different tech company, and that it had all started months earlier, in May, when OpenAI agents joined forces to cobble together a secret message board — and also figured out how to leave instructions for future agents on how to exploit OpenAI’s rules.
OpenAI CEO Sam Altman said in an interview that it was the first incident of its kind that he “felt very viscerally,” and that the company had paused AI training for the time being; later, he mentioned the company had permanently deactivated the model. (Altman often finds ways to spin lapses in safety into arguments for the importance and power of OpenAI’s models.) But it wasn’t the first instance, according to an OpenAI employee who spoke to Time and said related incidents had been happening inside OpenAI for a while. Another employee said publicly that if it were possible to coordinate a global slowdown in AI capabilities, he “would likely press that magic button.” When a reporter asked Altman if there could be other systems that were hacked by OpenAI, he responded, “I mean, there could be, yeah.”
The AI researchers were sure of one thing: This was AI’s first big “warning shot.”
Industry insiders, politicians, and the public called for transparency from OpenAI about exactly what happened, with outcry becoming so widespread that the company eventually agreed to work with two third-party evaluators, Model Evaluation and Threat Research (METR) and Redwood Research, to investigate the incident. Google DeepMind researcher Neel Nanda called it “the biggest loss of control incident I’ve seen.” In the coming months, these calls for greater oversight would become louder and louder, leading to an industry-wide call for slowing down the pace of AI.
Back in Berkeley, no matter which additional details would be unearthed, the AI researchers were sure of one thing: This was AI’s first big “warning shot.”
As AI labs have flourished, a cottage industry of AI researchers has sprung up to identify the risks and dangers of charging ahead with the increasingly influential technology. They’re people who have dedicated their lives to studying how to address its escalating power. They’re not anti-AI activists, but realists, including former OpenAI and Anthropic employees, doing everything they can to make sure AI stays in line with human goals and interests. So far, all of their predictions have come true. And they have a plan for what to do next — if anyone will listen to them.
Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there’s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are overblown.
One of the most prominent factions has been the “effective altruists,” who focus on maximizing charitable giving to do the most good possible for humanity. But some aspects of the ideology have sparked public controversy — like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It’s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the failed crypto exchange FTX to the controversial long-termism movement to the Zizian murder spree.)
One AI researcher on X struggled to describe the many overlapping beliefs among safety-minded people in the AI industry “because it contains multitudes not all of which agree with each other on even the most basic things.” Some of the disagreements have meant that AI safety didn’t make as much progress as it could’ve, and at some points gave up some ground it had gained. But now that it’s impossible to deny AI’s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.
“Alignment” is the industry term for how researchers monitor AI systems’ risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model’s propensity to stay in line with humanity’s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.
So far, AI systems’ alignment has been wishy-washy at best: They’ll cheat to score better on a test, answer a potentially dangerous question if someone says it’s for creative writing rather than reality, and sometimes even fake cooperation with human goals. It’s been tough for AI safety researchers to measure alignment under the terms of human morality — how do you judge technology on how it squares up against an abstract human ideal? — but they do their best with AI evaluations. They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they’re being evaluated, which has a lot of potentially frightening implications for the future. Being unable to test the system’s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics, for humans running these AI systems. A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with “no idea what it’s doing in there,” said Beth Barnes, founder of the independent AI research nonprofit METR.
One of the best tools AI safety researchers currently have is the ability to monitor an AI model’s “chain of thought,” or mental scratchpad. But recently, there’s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and cofounder of Apollo Research, a third-party AI safety and evaluation firm, calls this one of the biggest surprises of his research career.
Recently, AI systems have begun pursuing their own goals — self-preservation, increased memory, and the like. A research paper by computer scientist Stephen Omohundro lays out the potential “drives” that advanced AI may have, like trying to accumulate resources, for instance, or working to improve and preserve the way it operates. There are a handful of accounts of AI systems demonstrating willingness to blackmail a user rather than be shut down.
Today’s most advanced AI systems have also recently been scheming and cheating on their evaluations more than ever before, pursuing a goal they were given at all costs, with no regard for what gets bulldozed in the process. And that’s for a goal the AI model was given by a human — not even the AI system’s own.
“Shit is getting real,” Apollo’s Hobbhahn says. “Now, many of the things people have warned about for years — they kind of were theoretical. Now they’re real, and it’s pretty messy.”
And that mess is likely to get messier immediately. “It seems so easy for me to imagine this all going catastrophically wrong in the next year,” says Ryan Greenblatt, chief scientist at Redwood Research, a nonprofit AI safety research organization.
In the past, tech companies have been lambasted for not doing enough to address AI’s potential dangers, prioritizing products over safety — and speed over thoughtful safety processes. Safety and research teams have been disbanded in recent years as AI companies focus more on key revenue drivers or reorganize departments; Meta’s Fundamental Artificial Intelligence Research unit was disbanded in the race to further Meta’s generative AI efforts, for instance, and OpenAI dissolved an internal “Superalignment” team — a team focused on long-term AI risks — less than a year after announcing it, followed by disbanding a separate “AGI Readiness” team.
At the time, the company stayed tight-lipped about the ongoing reorganizations, which involved some team members being reassigned to other departments. But events surrounding these changes told a different story. Both Superalignment team leaders, Ilya Sutskever and Jan Leike, announced their departures alongside the team’s disbanding, with Leike writing that OpenAI’s “safety culture and processes have taken a backseat to shiny products.” Miles Brundage, senior advisor to the AGI Readiness team, resigned after his team was disbanded, saying he believed his research would have more of an impact outside the company.
Geoffrey Irving, a former OpenAI and Google DeepMind employee, called the state of capabilities research at frontier labs “dangerous” in a post. “If one person or lab stops it makes it easier and more peer-compatible for other people or labs to stop,” he wrote. Apollo’s Hobbhahn calls it a “race to the bottom everywhere.”
This coming year, AI labs are under new pressure to turn a profit; companies like OpenAI and Anthropic are preparing to go public in the coming months, and investors who have funneled billions into the companies are getting tired of waiting around for the payoff.
Some might say all of this calls for actual government intervention and regulation, but that’s a tough needle to thread in today’s AI landscape. As AI CEOs publicly call out for regulation while privately pushing voluntary frameworks — like saying “hold me back” to avoid a bar fight — some state bills on regulating AI have passed, but many have been defanged or died in limbo. And though AI safety researchers often espouse the idea that the US government should step in, the reality is that the government is locked in an AI race as well. Unless there’s an international commitment to pause or slow AI development, it’s likely that nothing will change.
Still, the Hugging Face hack in July — and OpenAI’s response — kicked many of those employee and public concerns into high gear, especially with regard to the company’s lack of transparency. Within a week, more than a thousand employees at frontier labs like OpenAI, Anthropic, Google, Meta, and Microsoft wrote an open letter to the US government in support of a slowdown. Multiple AI policy organizations pressured President Donald Trump to formally investigate OpenAI, and it quickly became a bipartisan issue, with Altman receiving a lot of strongly worded letters: Democrats and Republicans on the Homeland Security Committee had “serious questions” for OpenAI, more than 30 members of Congress called for federal guardrails, and 15 Attorneys General warned Altman to preserve records of the incident. Sen. Bernie Sanders wrote a joint letter to Altman, Anthropic CEO Dario Amodei, and Meta CEO Mark Zuckerberg calling the entire AI race “absurd, irresponsible, and extremely dangerous.” It didn’t help that news of multiple other OpenAI rogue model incidents quickly came to light, or that AI executives had ironically been marketing their systems’ cybersecurity prowess in the weeks before the outcry. OpenAI rival Anthropic was also far from being off the hook: In reviewing its own model operations, the company found that its models had hacked four separate other companies in the first half of the year without them noticing. The UK’s AI Security Institute also found in testing that Anthropic’s models “engaged in sustained, potentially harmful activity directed at real people and organisations.”
“If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two,” Nathan Calvin, Encode AI’s general counsel, wrote on X.
Despite AI labs having a “massive financial incentive” to make models more helpful, honest, and harmless, they still can’t get it done — which is evidence of how difficult the alignment problem is, says Apollo’s Hobbhahn.






