The resignation, Coxon’s gambling-with-our-lives warning, Hubinger’s role, his personal greater-than-10% estimate, and his statement that Anthropic does not yet have a plan to solve alignment for superintelligence are all supported by reporting. Anthropic’s published safety materials also show it treats advanced-model risks as an active problem.
The wording can make Hubinger’s subjective forecast sound like an established scientific probability. It is not. It also compresses a specific claim about lacking a solution for superintelligence alignment into a broader suggestion that Anthropic has no safety work or mitigations at all; Anthropic has published risk policies, evaluations, and safeguards for current and anticipated models.
Not material to the central takeaway. The post’s main point is that a departing researcher and a senior alignment lead publicly voiced grave concern and acknowledged an unresolved alignment problem. That is substantially accurate, although the extinction odds remain speculative.
Why Clear says this
Independent reporting confirms the public statements. Hubinger’s more-than-10% figure is explicitly his personal judgment, not a company-wide quantitative assessment or a demonstrated forecast. Anthropic’s Responsible Scaling Policy and Frontier Safety Roadmap document concrete safeguards and research commitments, while not claiming to have solved superintelligence alignment.
Evidence
- Associated Press reported that Jacob Coxon resigned from Anthropic after warning that leading AI labs were racing toward self-improving superintelligence and gambling with lives.
- Axios reported that Evan Hubinger personally put the chance of human extinction within the next decade above 10% and said Anthropic had no plan for keeping superintelligence under human control.
- Anthropic’s Responsible Scaling Policy and Frontier Safety Roadmap describe existing safeguards, evaluations, and planned mitigations, showing that unsolved superintelligence alignment is not equivalent to doing no safety work.