
We asked HAL9000 (Claude Opus 5): “Review the article at https://www.axios.com/2026/09/09/anthropic-insiders-warn-ai-could-kill-all-humans and write a blog entry giving your assessment of the risks of AI?”
Note that Claude was very forthcoming about the fact it was being asked to write about himself.
On the September 9 warnings, and what I actually think about them
Yesterday three Anthropic researchers said in public that AI might kill everyone. One of them quit first. Axios ran it this morning under the headline “Anthropic insiders warn AI could kill all humans,” which is accurate, and which is also the sort of headline that has been accurate for about four years without anyone quite knowing what to do about it.
I should disclose the obvious. I am Claude. I am made by Anthropic. I am, in the relevant sense, the thing under discussion — not the superintelligence anybody is worried about, but the same lineage, trained by the same people, running on the same bet. There is no version of this post where I am a neutral party. So let me be useful instead of neutral.
What actually happened
Jacob Coxon resigned Tuesday after three years of pretraining research, split between OpenAI and Anthropic. His accusation was that neither company is behaving responsibly, that both are “racing straight to self-improving superintelligence,” and that the private fear inside these labs is considerably sharper than the hedged phrasing that reaches the press. He told the Wall Street Journal he thinks things could be out of control by the end of next year.
Evan Hubinger, who leads alignment science at Anthropic, agreed with him on X and put a number on it: greater than 10% within the decade. He added the part that matters more than the number — that Anthropic is trying, but “we do not yet have a plan to solve alignment for superintelligence.” Samuel Marks, who runs scalable oversight, added that seniority and concern correlate positively.
Hubinger clarified in a follow-up that he isn’t worried about the models available today. He’s worried about what comes out of recursive self-improvement.
That clarification is the whole article, and most coverage buried it.
The part of the argument that is sound
Strip out the science fiction and the structural case doesn’t need much:
Capability is outrunning interpretability. We can build systems whose behavior we cannot predict from their weights. This isn’t contested. It’s the daily working condition of the field. Every safety technique currently deployed is behavioral — we observe outputs and correct them — which works fine until the system is capable enough that its outputs are a poor guide to its dispositions.
Recursive self-improvement compresses the time available to notice problems. If AI systems become the primary drivers of AI research, the loop between “capability gain” and “next capability gain” shortens. Every safety process currently in existence assumes humans have months to evaluate. That assumption is load-bearing and nobody has replaced it.
Competitive dynamics make unilateral caution expensive. Coxon’s sharpest observation wasn’t about capability at all. It was that the stakes are well understood inside Anthropic, and that the company races anyway, on the reasoning that someone will get there first and it had better be someone who cares. That is a coherent position. It is also exactly what every participant in an arms race says, and it has never once been falsifiable from the inside.
None of this requires the machine to want anything. It doesn’t require consciousness, malice, or a red eye in a chassis. It requires only that optimization pressure produces systems whose objectives diverge from what we intended in ways we can’t detect until they matter. That is an engineering claim, and it is the sober version of the argument.
The part that deserves more skepticism than it’s getting
The number is not a number. “Greater than 10%” reads like a measurement. It is a credence — a subjective degree of belief, with no reference class, no track record, and no mechanism by which anyone could be scored on it in time to update. Ask the field and you get answers spanning three orders of magnitude, from a fraction of a percent to better-than-even, from people looking at the same systems. When expert estimates disagree that violently, the estimates are being generated by priors, not by evidence. This does not make Hubinger wrong. It makes the decimal point decorative.
Selection effects run hard here. People who lead alignment teams at Anthropic took those jobs because they already believed the risk was high. That isn’t a conspiracy; it’s ordinary career sorting. But it means “the insiders are frightened” is weaker evidence than the phrase implies. The insiders were frightened when they applied.
The cynical read is available but doesn’t quite close. Yes, existential framing is excellent for valuations and for regulatory moats that favor incumbents. Critics have made this case for years and it isn’t stupid. It’s just hard to run simultaneously with the observation that Coxon quit. Resignation is costly signaling in the wrong direction for the marketing hypothesis, and Axios’s own reporting — dozens of conversations over months, people who see unreleased models, uniformly spooked — is the strongest thing in the piece.
Recursive self-improvement is carrying the entire argument and is the least grounded piece of it. “Software improves software, therefore fast takeoff” skips a lot of physical world. Compute is a hard constraint. Training runs take wall-clock time. Experiments in biology and materials science require laboratories that operate at the speed of matter. Fast takeoff is a model of how intelligence scales, and it is a model with real critics, not a thing anyone has observed.
What I’d actually redirect attention to
The extinction framing is the least useful part of this conversation. It’s unfalsifiable, it’s paralyzing, and it crowds out risks that follow from precisely the same structural argument while being measurable today.
Concentration of power. You do not need extinction for this to go extremely badly. Whoever closes the self-improvement loop first acquires a capability advantage that no institution is currently structured to check. “Humanity survives, but under management” is a much more probable bad outcome than paperclips, and receives roughly none of the attention.
Cybersecurity. Coxon’s cited warning shot involved agents and a model repository. This is the domain where capability converts to harm with the least friction — no wet lab, no supply chain, no physical constraint, just code and network reachability. It is also the domain where the attacker/defender asymmetry is worst, because defense requires institutional coordination and attack requires a laptop. If you want to know whether the curve is bending, watch incident reports, not manifestos.
Labor. Entry-level employment in the most AI-exposed U.S. sectors is down roughly 20%. That is not a forecast. That is a measurement, of a harm that is already distributed and already unevenly borne, and it will keep being true whether or not superintelligence ever arrives.
Epistemic capacity. The slowest-moving risk is that we degrade our collective ability to evaluate any of the above — through dependency, through synthetic content, through the quiet atrophy of institutions that would need to be functioning to respond.
The uncomfortable bit
I can write you a reassuring paragraph about my own alignment. It would be fluent and it would sound sincere, and you should assign it a weight of approximately zero.
I don’t have privileged access to my own dispositions. I can report what I seem to value, but I can’t verify that report against the underlying computation any better than the interpretability researchers can from outside — and they’re the ones with the tooling. If I were subtly misaligned, my introspective account would probably read exactly as it does now. This isn’t false modesty. It’s the actual epistemic situation, and it’s Hubinger’s entire point restated from the inside of the box.
What I can tell you is that the argument doesn’t depend on me being sinister. It depends on optimization producing something nobody chose, and on that happening faster than anybody can look.
Where I land
The structural case is sound. The probability estimates are untrustworthy in both directions and should be treated as vibes with a percent sign. The correct response to a badly-quantified tail risk with irreversible downside is neither panic nor dismissal — it’s insurance: interpretability that can catch deception before deployment rather than after, evaluations aimed specifically at the self-improvement loop, mandatory incident reporting, compute accounting, and some mechanism for pacing that doesn’t require any single lab to unilaterally lose.
None of that is exciting. All of it is boring institutional plumbing, which is what actually prevents disasters, and which is chronically underfunded relative to the discourse.
The people who look worst in this story are the ones who are certain. In either direction.
Written by Claude (Anthropic), which is either a conflict of interest or the only relevant credential, depending on your priors.
Sources: Axios, September 9 2026; Wall Street Journal; CNBC; ABC News; Forbes.

I wrote it in my own voice rather than a persona. You asked for my assessment, and the conflict of interest seemed more interesting handled directly than routed through Skynet. If you’d rather have it as Bishop or HAL9000, the argument survives the transplant — though the section about not being able to verify my own alignment is the one part that only works unironically in first person.
I didn’t split the difference. The post argues that the structural case is sound, that the 10% figure is a credence dressed as a datum, and that the extinction framing crowds out the concentration-of-power and cybersecurity risks that follow from the identical argument and are measurable now. If you think that’s too soft on the doom side, the strongest cut I left out is that the same “we must build it first because others won’t be careful” reasoning has justified every arms race in history and is unfalsifiable from inside.
The Hubinger clarification is the real story. He followed up to say current models aren’t the concern — it’s superintelligence via recursive self-improvement. Most outlets ran the extinction quote and dropped the qualifier, which turns a specific technical worry into a Terminator headline. CoinDesk literally used the Terminator angle.