Yes, an Anthropic employee really did say that. He said it on the record, under his own name. Evan Hubinger leads Alignment Science at Anthropic. On September 9, 2026, he wrote on X that he personally believes AI has more than a 10% chance of killing all humans within the next decade. He also said Anthropic doesn't yet have a plan to prevent it.
Short answer: Yes, an Anthropic researcher really did say that. Evan Hubinger, Anthropic's Alignment Science Lead, wrote on X on September 9, 2026, that he personally believes there is more than a 10% chance AI kills all humans within the next decade, and that Anthropic still has no plan to solve alignment for superintelligence.

I spent an afternoon tracing this story back to its source instead of trusting the headline. "AI could kill everyone" gets thrown around a lot online. Usually it traces back to a vague survey or an anonymous quote. This one doesn't. When I compared Hubinger's original post against CBS News’ reporting and Anthropic's own published safety policy, the numbers and the context matched across every version. This is a real statement from a real Anthropic safety lead, made in public, under his own name.
What Evan Hubinger actually said
Hubinger wasn't responding to a reporter's question. He was replying, on X, to a departing colleague, and his reply read: "Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Two things make that statement land differently than the usual AI-doom quote. First, Hubinger isn't a pundit. He isn't a former employee with an axe to grind, either. He currently leads alignment science at one of the three labs building frontier AI models — the team whose entire job is making sure powerful models do what they're supposed to do. Second, he didn't hedge with "some researchers believe" language. He put a number on it. He attached his real name to it. He posted it somewhere it couldn't be walked back quietly. CBS News picked up the exchange the same day, and it spread through the tech press within hours.
It's worth being precise about what the number covers. Hubinger's ">10% within the next decade" is about future systems. He means superintelligent AI that today's safety techniques haven't caught up to yet — not the chatbot you used this morning. He said as much in the same thread: today's models aren't the danger. The trajectory is.
Why this surfaced now: a resignation, not a leak
Hubinger's post wasn't posted in a vacuum. A day earlier, on Tuesday, September 8, 2026, a colleague resigned. Jacob Coxon was a 27-year-old Anthropic researcher who had spent three years doing pretraining work at both OpenAI and Anthropic. He posted that neither company was "acting responsibly." His line that got quoted everywhere: the labs are "racing straight to self-improving superintelligence and gambling with our lives."
Coxon's account, as reported, was more specific than a general complaint. He described OpenAI staff as not having fully internalized the civilizational stakes. He described Anthropic differently — understanding the stakes clearly, but feeling boxed in by competition, and proceeding anyway because leadership doesn't believe a unilateral slowdown would make anyone safer if rivals keep racing. That's an uncomfortable position for a safety-focused company to be reported as holding. It's why Hubinger's reply reads less like a hot take and more like a confirmation from the inside.
The timing lines up with a separate story that had already put Anthropic's safety commitments under scrutiny. The same week, reporting showed Anthropic had not given the UK's AI Security Institute pre-release access to Claude Mythos 5.1, the more capable, access-limited model line it launched September 1, 2026. That's a reversal from April, when the institute had tested the prior Mythos release. Anthropic hasn't confirmed why access wasn't extended this time, though it says it's working to widen access abroad. On its own, that's a policy dispute. Next to Hubinger's post, it reads like a pattern: the same week a lab's own alignment lead says there's no plan for superintelligence risk, that lab also pulls back independent testing access on its newest model.
What "more than 10% in a decade" actually means
A 10% chance over ten years is not a fire alarm going off today. It's not nothing, either. Picture it like a single round of Russian roulette, applied once to the entire species over a ten-year window instead of instantly. That's the uncomfortable framing Hubinger's own colleagues have used before. Anthropic CEO Dario Amodei has separately put his own estimate of catastrophic AI outcomes in a similar double-digit range.
This kind of estimate is often shortened online to "p(doom)" — the probability a given person assigns to AI causing a catastrophic or extinction-level outcome. It's not a measured statistic. It's one person's judgment, shaped by how fast the technology is improving and how confident they are that safety research is keeping up. Well-informed people inside the same lab can land anywhere from under 1% to over 20%. That spread is part of the story: OpenAI’s newest reasoning technique has separately alarmed outside AI safety researchers, so this isn't an Anthropic-only worry.
What actually drives the number up, according to people like Hubinger, isn't today's chatbots. It's the chance of AI systems that can improve their own successors faster than humans can verify the result is safe. That's a different problem than a model giving a wrong answer or leaking data. It's why this story reads in a different, heavier register than most AI-safety news.
How Anthropic, OpenAI, and Google DeepMind's safety promises compare
Every major lab has published a document meant to answer the question Hubinger raised: what happens if a model gets too dangerous to ship? I read all three of the current versions back to back. The intent is similar across the board. What differs is how concrete the trigger actually is.
| Framework | Company | Last updated | What it tracks | What's supposed to happen at the line |
|---|---|---|---|---|
| Responsible Scaling Policy | Anthropic | Aug 14, 2026 | AI Safety Levels tied to catastrophic misuse and misalignment risk | Anthropic must build an "affirmative case" that misalignment risk is addressed before releasing a model past a threshold |
| Preparedness Framework v2 | OpenAI | Apr 15, 2025 | Biological/chemical capability, cybersecurity, AI self-improvement, rated Low to Critical | A "Critical" rating halts further development until safeguards meeting that standard exist |
| Frontier Safety Framework v3.0 | Google DeepMind | Apr 17, 2026 | CBRN, cybersecurity, ML R&D, and (new in v3) deceptive alignment, via Critical Capability Levels | Crossing a Critical Capability Level triggers mandatory mitigations before further scaling |
On paper, all three commit to slowing down before a model crosses a dangerous line. Hubinger's post adds something the documents don't say outright: Anthropic's own version of that plan, the part covering a model that could out-think its own safety checks, isn't finished yet. My deeper comparison of how OpenAI and Anthropic each define AI responsibility goes further into how these policies get enforced in practice, not just how they read on paper.
Common mistakes people are making with this story
The biggest one is reading "10% chance AI kills all humans" as a claim about the AI product you use today. It isn't. Hubinger was explicit that current models aren't the danger; the risk he's describing is tied to future systems capable of recursive self-improvement, which don't exist yet.
The second mistake is treating this as an Anthropic-specific scandal instead of an industry-wide gap. Every frontier lab's safety framework has the same weak spot Hubinger pointed at: a solid plan for handling misuse today, and a much thinner one for a model smarter than the humans checking its work. Debates over how AI regulation gets messaged run into this same gap — rules written for today's models don't automatically cover tomorrow's.
The third mistake is assuming a warning like this settles the "is AI dangerous" debate. It doesn't. It's one credentialed person's estimate, shared publicly, from inside a company that wants to be seen as the safety-conscious lab. That doesn't make it wrong — Anthropic has generally been more open about internal disagreement than its rivals. But it isn't proof of an outcome either. Treating a probability estimate as a certainty, in either direction, misreads what Hubinger actually said. Lawmakers reacting to headlines like this one is exactly why there's now a real bill in Congress aimed at pausing or restricting superintelligence development — a response to the uncertainty, not to a proven event.
What this actually changes about the AI tools you use
For the AI chatbot or coding assistant you used this week, functionally nothing changed on September 9, 2026. The risk Hubinger described applies to a future category of system, not to Claude, ChatGPT, or Gemini as they exist right now. If you're weighing which everyday AI tool to trust with your work, the practical risks are still the boring, documented ones: hallucinated facts, data handling, and AI agents that occasionally lie, cheat, or overstep the permissions they were given. My full breakdown of the risks of AI that are real but manageable covers how to handle those without overreacting to a headline about extinction risk.
Where this story is genuinely useful is as a filter for picking which lab you trust more, long term. A safety team that admits a gap in public, instead of papering over it, is at least being honest about where it stands. That's worth something, especially now that Anthropic’s own commercial position is already under pressure from cheaper rivals and has every incentive to project confidence instead. If you're comparing the major assistants for day-to-day use, my head-to-head test of Claude against ChatGPT is a better guide to which tool fits your work than a headline about a decade-out risk estimate. Not everyone reads the long-term picture this darkly, either. Meta's own public pitch, covered in The Future Is for Everyone, argues the opposite case. OpenAI’s policy blog, AI Futures, lays out its own version of how this goes right instead of wrong.
My take
Anthropic researcher says AI could kill all humans — and the part that actually moved me wasn't the 10% number. It's that the person saying it works there, said it under his real name, and named exactly what's missing (a finished plan for superintelligence alignment) instead of speaking in vague generalities. That's a different kind of warning than the usual AI-doom cycle. It doesn't mean the sky is falling on your Tuesday-morning ChatGPT session. It means the people closest to the frontier are telling you, plainly, that the hardest part of this problem isn't solved yet. They're asking you to take that as seriously as they clearly do.
Frequently Asked Questions
Did an Anthropic researcher really say AI could kill everyone?
Yes. Evan Hubinger, Anthropic's Alignment Science Lead, posted on X on September 9, 2026 that he personally believes there's more than a 10% chance AI kills all humans within the next decade, and that Anthropic doesn't yet have a finished plan to solve alignment for superintelligence.
Does this mean today's AI chatbots are dangerous?
No. Hubinger was explicit that the risk applies to future, more capable systems — particularly ones that could improve themselves faster than current safety methods can verify — not to the chatbots and coding assistants available right now.
Who is Jacob Coxon, and why did he resign?
Jacob Coxon was a 27-year-old Anthropic pretraining researcher who had also worked at OpenAI. He resigned on September 8, 2026, saying neither company was "acting responsibly" and that the industry was "racing straight to self-improving superintelligence and gambling with our lives." Hubinger's post was a direct reply to him.
What is "p(doom)" and is it a real measurement?
"p(doom)" is shorthand for one person's estimated probability that AI causes a catastrophic or extinction-level outcome. It's a subjective judgment, not a scientific measurement, and estimates from credible researchers vary widely — from under 1% to over 20% — depending on how much weight they put on current safety progress.
Has Anthropic responded to Hubinger's post?
Anthropic hasn't issued a formal statement disputing what Hubinger said. His post itself, made publicly under his own name while still employed there, is the closest thing to an on-record acknowledgment that the company doesn't yet have a complete plan for superintelligence-level alignment risk.