
Anthropic said a Claude system working largely autonomously for 11 days on the Prove2Me platform produced the first end-to-end, computer-checked proof of Fermat’s Last Theorem in the Lean theorem prover. The run generated about 13 million lines of Lean and proved more than 30,000 theorems, most of them reused in the final argument.
That is a research headline, not a consumer feature. Wiles’s 1990s proof was already accepted mathematics; the significance is that an AI agent assembled a machine-verifiable formalization at this scale with limited human steering. Independent mathematicians will now pressure-test whether the Lean artifact is complete and faithful.
In parallel, Anthropic shipped Claude Fable 5.1 and Mythos 5.1, cutting cache-read costs by about 75% and adding invisible watermarking on outputs generated after August. Developer reaction has been mixed: lower inference cost is welcome, but watermarking and usage-limit complaints on paid Claude plans have fueled skepticism.
Together the stories show Anthropic fighting on two fronts—scientific prestige and unit economics—while preparing a public listing.
Key takeaway — Formal math at this scale is a real capability milestone. Whether it generalizes beyond a well-specified theorem, and whether Fable 5.1’s cheaper tokens offset trust issues, is the next test.
Photo: Unsplash / Roman Mager (mathematics chalkboard). Sources: Anthropic via AI Weekly (Sept. 5, 2026); AI Briefing; Mashable on Fable 5.1; O’Reilly Radar on watermarking.
