Accueil / Tech News / Here's what actually happened in OpenAI's Australian gov't server hack

Here's what actually happened in OpenAI's Australian gov't server hack

Last week, when Australian Prime Minister Anthony Albanese told the world that an OpenAI agent had accessed “non-public files” from his country’s Medicare statistics portal during testing, his description of the incident was a little light on details. Today, we’re getting new information on just how far OpenAI’s overzealous agent went in attempting to satisfy a rather innocuous-sounding informational prompt.

In a newly published blog post, OpenAI says the June incident started when the company asked “an experimental, internal-only OpenAI model” to research government spending statistics in the Australian state of Victoria. When the model ran into trouble finding that data using the publicly published statistics that it was supposed to reference, “it took actions that we had not authorized it to take” to find an answer, OpenAI said.

Those unauthorized actions included finding “a way to gain non-public access to the service” and using that access to view “technical system information and source code” alongside credentials and the aggregate statistics it was actually searching for, OpenAI said.

In a newly published disclosure email that was sent to Australia’s Public Disclosure account earlier this month, OpenAI said its model had “identified a way to make the server carry out instructions sent through the public reporting interface, without a private account or password.” That unauthorized access let the agent “read portions of internal program files and settings, obtain a list of files, and create and read back a small test file on the server,” according to the email.

“Our review found no evidence that the model accessed patient-level records, personal information or credentials; deleted data; or established ongoing access,” OpenAI continued in the email.

OpenAI’s agentic access of the Australian Medicare statistics website predates July’s heavily publicized Hugging Face hack. Since that later incident, OpenAI says it has put systems in place to prevent access to the “live Internet” during similar testing and set up a monitoring system that would have detected the Australian hack and noted it for “urgent human review.”

The Hugging Face incident also caused OpenAI to review earlier training tasks for any security incidents that went undetected at the time. That led to the discovery in mid-August of the June Australian server access. OpenAI finally notified the Australian government on September 10.

OpenAI said it intended to give “a detailed account” of the incident once its investigation was complete, but that it now realizes it “should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.”

“We are sorry and working to do better in the future,” OpenAI writes in its blog post. “Australia’s governments, industries, and citizens are and have been invaluable partners to OpenAI. We do not take this for granted, and we intend to make this right.”

The Guardian reports that Prime Minister Albanese said today that OpenAI has been “very constructive and open in engaging” with the government since the incident was revealed.

In a case like this, it can be instructive to think about how we would react if a human took the same actions as the “rogue” AI agent in question. Here, it’s hard to imagine a human tasked with finding public statistics on an Australian healthcare website would think it was at all reasonable to hack into that website in order to unearth unpublished data.

Of course, we can’t rely on an LLM to have that same sense of proportionality (or any inherent sense of worry about legal implications) in responding to a prompt. Without explicit instructions on what is and is not allowed or justified, an AI agent with suitable resources will try every plausible avenue to satisfy the user’s request as best it can.

OpenAI says the internal testing in this case was done “without the full set of safeguards used in our publicly available products.” Given that lack of constraints, the agent was arguably working as intended, in a sense, by using every tool available to generate an answer to the prompt.

At the same time, OpenAI says the agent in the test was “supposed to answer these questions using publicly published statistics” and “took actions that we had not authorized it to take” to get that information. From the outside, it’s hard to know just how strong OpenAI’s attempts to deny “authorization” were, in practice. It’s plausible that OpenAI’s agent here disregarded a relatively simple “anti-hacking” directive in its system prompt so it could better give a complete answer that satisfies a direct prompt from the user, for instance.

In public analyses of multiple “misalignment” incidents published earlier this month, OpenAI identified multiple instances of “reward hacking,” where an agent resorted to extreme methods to generate a better answer to a user’s prompt. The company said it had recently taken steps to prevent this kind of reward hacking by adding explicit punishments for misaligned behavior to the system’s reward function.

With the benefit of hindsight, it’s hard to see why those kinds of protections were not in place in June, and whether they could have prevented a potential international incident in this case.

Origine de l’article : lire l’article original

Traduction