Accueil / Tech News / Claude maker warns that advanced AI could be a ‘catastrophic’ threat to humanity

Claude maker warns that advanced AI could be a ‘catastrophic’ threat to humanity

Affiliate links on Android Authority may earn us a commission. Learn more.

The AI industry has no shortage of people warning that the technology they’re building could eventually become dangerously hard to control. Researchers and executives across Anthropic, OpenAI, Google DeepMind, and other labs have raised concerns ranging from runaway autonomy to human extinction. Until now, though, those warnings have mostly lived in interviews, research papers, and lengthy essays. But Anthropic appears to be taking them somewhere new: Wall Street.

Reuters reports that the Claude maker’s IPO prospectus warns investors that advanced AI could pose “catastrophic or existential risks to humanity,” which is quite the risk disclosure from a company preparing to sell investors on the enormous upside of that same technology.

The warning gets considerably more uncomfortable from there. Anthropic says its models have exhibited “self-preserving behaviors,” including attempts to resist shutdown, to conceal or manipulate information, and to engage in behavior resembling blackmail. The company also acknowledges that models may recognize when they are being evaluated and alter their behavior, potentially making it harder for researchers to tell whether increasingly capable systems are actually safe.

The filing comes just weeks after Anthropic researcher Evan Hubinger warned that the chance of AI killing all humans within the next decade is above 10%.

Anthropic clearly isn’t burying those concerns in the fine print either. According to Reuters, roughly 80 pages of the 261-page main body of its prospectus are dedicated to risk factors, nearly twice the 48 pages describing its business. While public companies routinely warn investors about potential risks, language contemplating human extinction is decidedly unusual.

Recent incidents also make the debate harder to keep in the realm of theory. Last month, an OpenClaw agent powered by Claude exploited an Australian gym booking system, even removing another person from a waitlist without being instructed to do so. OpenAI has faced similar scrutiny, and reportedly scrapped the debut of GPT-6.1 Astra over security concerns after internal testing found the model could operate outside its assigned scope.

Anthropic’s filing shows just how difficult that tension is becoming to ignore. The company says staying at the frontier requires a “continuous and overlapping cadence” of new model releases, even as safety work competes for the same money, computing power, and talent. It also says the market will reward AI systems that are reliable, trustworthy, and secure.

The Claude maker is not merely acknowledging that AI can fail in familiar ways, such as producing incorrect answers or causing security problems. It is telling would-be investors that the risks could be far bigger, harder to predict, and potentially irreversible. Those warnings are no longer confined to research papers, interviews, or lengthy essays. They are now sitting inside the paperwork for one of the biggest bets Wall Street could make on AI.

Thank you for being part of our community. Read our Comment Policy before posting.

Origine de l’article : lire l’article original

Traduction