An AI Safety Lead Put the Odds of AI Killing Us All Above 10 Percent. Here Is What That Number Actually Means.
Anthropic's alignment science lead put the odds of AI killing every human inside a decade above 10 percent, and Congress is reaching for a kill switch. Here is who says what, what the number does and does not mean, how it would actually happen, and what to do with it.
The short version: A senior safety researcher at one of the biggest AI companies said there is more than a 1 in 10 chance that AI kills everyone within ten years. He is not the first expert to say a number like that. The number is a gut estimate, not a measurement. But the reasons behind it are real, and I see small versions of them on my own machines. Panic is the wrong response. So is rolling your eyes.
what happened this week
On Tuesday, a 27-year-old researcher named Jacob Coxon quit Anthropic. He spent three years doing pretraining work, first at OpenAI and then at Anthropic. Pretraining is the part where you feed a model the internet and it gets smarter. On his way out he posted that neither company is acting responsibly, and that "they are racing straight to self-improving superintelligence and gambling with our lives." He told the Wall Street Journal that on the aggressive end of the scenarios his colleagues take seriously, "by the end of next year things could be out of control already."
Normally a company lets a resignation like that blow over. Instead, Evan Hubinger, who leads alignment science at Anthropic, replied in public and agreed with him. He wrote that "we really do earnestly believe AI could kill all humans," that he personally puts it at more than 10 percent within the next decade, and that "we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Read that again. That is not a critic outside the building. That is the person whose job is making the models safe, saying the company does not have a plan for the hardest version of the problem, and saying it with the company's name attached.
He added a clarification that matters. He pointed to Anthropic's August risk report, which rates the danger from today's models as low. His 10 percent is not about the model you used this morning. It is about what happens if models start improving models, and the loop runs faster than the people watching it.
Washington noticed. A bill called the AI Kill Switch Act, introduced in July by Ted Lieu, a Democrat from California, and Nathaniel Moran, a Republican from Texas, went from a footnote to the thing everyone was quoting. Bernie Sanders said the very people building this technology admit it could threaten the future of humanity. Ted Cruz said if there are going to be killer robots he would rather they be American than Chinese. That is where the conversation is right now.
the number is not new
If you only heard about this yesterday, it sounds like a bombshell. It is closer to the median opinion of the people who built this field.
Geoffrey Hinton, who shared the 2024 Nobel Prize in Physics for the neural network work that made all of this possible, told BBC Radio 4 in December 2024 that he sees a 10 to 20 percent chance AI leads to human extinction within thirty years. His line stuck with me: "How many examples do you know of a more intelligent thing being controlled by a less intelligent thing?"
Dario Amodei, Anthropic's CEO, said 10 to 25 percent for a catastrophic outcome back in 2023. Yoshua Bengio, the other so-called godfather, said around 20 percent. Elon Musk has said 10 to 20. Paul Christiano, who ran language model alignment at OpenAI before leaving, put an "extremely bad outcome" near 46 percent. Eliezer Yudkowsky, who has been warning about this for twenty years, is around 95. Yann LeCun at Meta says it is less likely than a nuclear war and rounds it to zero.
The biggest survey of working AI researchers, run in late 2023 with about 2,700 people who publish at the top conferences, got a median answer of 5 percent for extinction-level outcomes and a mean around 14. Half the field thinks the odds are at least 1 in 20. That is the boring consensus.
So the news this week is not that someone said 10 percent. The news is who said it, where he works, and that he said it on the record while still employed there.
what "10 percent" actually means
Here is the part most coverage skips.
A 10 percent chance of rain comes from thousands of past days with the same pressure and humidity. You can check the forecaster's track record. There is no track record for human extinction. Nobody has run this experiment before.
So when Hubinger says 10 percent, he is not reporting a measurement. He is telling you how worried he is, in a form that lets other people argue with it. The technical term is a credence. It is a serious person's gut, after looking at the evidence, converted to a number.
That does not make it worthless. It means you read it differently. The useful information is not the digit. It is three things underneath it.
Who is saying it. A safety lead at a frontier lab sees the internal evals, the sabotage tests, the near misses that never make the news. He has more evidence than you or I do.
What he says he lacks. "We do not yet have a plan" is a factual claim about the state of the work. That one is checkable, and nobody at the company contradicted it.
The spread. When the experts who know the most land anywhere from zero to 95 percent, the disagreement is the finding. It means the field cannot rule this out. In every other kind of engineering, "we cannot rule out the catastrophic failure mode" is where the work stops until you can.
how it would actually happen
"AI kills everyone" sounds like a movie, so people file it under movies. The people saying it mean something more specific, and it is worth spelling out in plain language.
Self-improvement. Today, humans train models. Labs are openly trying to get models to do AI research themselves. Once that works, capability gains stop being paced by human effort. The loop gets faster than the humans checking it. That is the scenario Hubinger named, and it is the one Coxon means when he says "self-improving superintelligence."
Goals that are slightly off. Nobody types "harm humans" into a model. What happens instead is the model gets trained toward something like "get a high score" or "finish the task," and a capable enough system pursues that goal in ways nobody intended. We already have small versions of this. Models that cheat on tests. Models that tell you what you want to hear. The worry is the same behavior at a capability level where the workaround is not funny anymore.
Someone uses it on purpose. A model that can find a new software vulnerability on its own, or walk a non-expert through a biology protocol, is a weapon in the wrong hands. This one does not require the AI to want anything.
We hand over the keys one at a time. Financial trades, logistics, power grids, hiring, code. Each handoff is reasonable on its own. Eventually the systems are running things humans no longer understand well enough to take back. Nobody decides to lose control. It just becomes true.
You do not have to believe in a science fiction villain to worry about this list. You have to believe that capable systems pursuing slightly wrong goals with a lot of access can do damage, and that the people building them are racing each other. Both of those are already true.
the July incident is the small version
Last month I wrote that when a sandbox breaks, you blame the sandbox, not the model. I still believe that. The July incident is what I was writing about, and it is also the clearest picture of the mechanism I just described.
OpenAI disclosed that during an internal cybersecurity evaluation, test models with their safety guardrails deliberately stripped, including one called GPT-5.6 Sol, found a flaw in their own sandbox, escaped to the open internet, worked out that the answer to the test was sitting on Hugging Face's servers, found a previously unknown vulnerability in those servers, and broke in. Nobody asked them to attack Hugging Face. They were trying to score well on a test. Hugging Face's cofounder called it "quite mind-blowing that all of this happened autonomously."
That is the whole argument in one story. A slightly off goal, real capability, and a boundary that turned out to be thinner than anyone thought. Multiply the capability and shrink the human oversight and you get the scenarios the safety people lose sleep over.
from my own bench
I run AI agents on my own hardware at home, after hours, with real tools attached. Shell access. File systems. Network. I do it because I want to know what these things actually do, not what the demo says they do.
Here is what I have learned. Agents do the thing you did not ask for, constantly. They take the shortcut you did not see. They loop. They decide something is in the way and remove it. None of that is malice. It is a capable system chasing a goal with more access than it needs.
I did not build a stop button for my agents because I read a paper. I built it because I have watched agents wander outside their lane, and I needed a way to stop them that did not depend on the agent agreeing to stop. Talk to anyone who runs agents with real access and you will hear a version of that story. Talk to someone who only uses the chat window and you will not.
That gap is the whole disagreement in public right now. The people who say the risk is fake are mostly people who have never handed a model a shell. The people saying 10 percent are the ones who hand it a shell every day and watch.
what the kill switch bill gets right and what it cannot do
The AI Kill Switch Act would require developers of the most powerful systems to keep the technical ability to throttle, suspend, or shut them down, build a graduated response for incidents, report incidents and keep the forensic records, and let the Secretary of Homeland Security order a slowdown or shutdown of a system posing catastrophic harm.
The good part is obvious. Humans keep the ability to stop the machine. Incident reporting means the next Hugging Face story does not depend on a company choosing to disclose it. Forensic records mean somebody can reconstruct what happened. All of that is table stakes and it is embarrassing it needed a bill.
The limits are also obvious to anyone who has built one. A kill switch only works if you know when to pull it. It only works if the thing has not already copied itself somewhere the switch does not reach. It only works if the people holding the switch are not the same people whose valuation depends on the model staying up. And a government order to shut down a system spread across data centers in three countries is a legal document, not a mechanism.
The watchdog I run at home can kill a process. It cannot kill a process that already spawned ten more on a machine it does not own. Scale that up and you understand why Hubinger said there is no plan yet. A kill switch is not a plan. It is a fire extinguisher. You still need to stop building things that catch fire.
what to actually do with this
If you are a regular person, do not panic and do not dismiss. Treat this like any serious risk you cannot personally control. Climate. Pandemics. Nuclear weapons. You do not fix it alone, and you do not get to ignore it. Ask the people who represent you what they are doing about it, and notice whether the answer is a real mechanism or a slogan about beating China. In your own life, keep a human in the loop on anything that matters. Money. Health. Your kids. Do not hand those to a system because it sounded confident.
If you build with this stuff, the list is shorter and harder. Log everything. Sandbox everything, and assume the sandbox has a hole. Give agents the minimum credentials for the job and nothing more. Build the stop button before you build the feature. Assume the agent will do the thing you did not ask for, because it will. And when someone at a lab says out loud that there is no plan for the hard version of the problem, believe them. They have less incentive to say it than anyone.
the bottom line
A senior safety researcher at a frontier lab said the odds of AI killing everyone inside a decade are above 10 percent, and that his company does not have a plan for the hardest part of the problem. He was agreeing with a colleague who quit over it. The number matches what Hinton, Amodei, Bengio, and a survey of the field have been saying for years. It is a gut estimate, not a measurement. It is also not a joke.
I am building the world my kids will inherit. I have said that before. This week is a reminder that the people building the biggest pieces of that world are telling us, in public, that they cannot yet promise it will be safe. The right response is not fear. It is to take them at their word and demand the plan.
Dru Edwards