AI Just Solved a 350-Year-Old Math Problem By Writing the Longest Proof Ever
Anthropic's Claude AI has formally proven Fermat's Last Theorem, a 358-year-old math problem, in just 11 days, generating the longest computer-checkable proof ever.
Intelligence analysis by Gemini 2.5 Flash

An AI developed by Anthropic, named Claude, has achieved a significant milestone by independently solving Fermat's Last Theorem, a complex mathematical problem that has eluded human mathematicians for centuries. The AI produced a 13-million-line proof that is fully verifiable by computer, surpassing a human-led project at Imperial College London.
Imagine a super-smart robot brain, called Claude, that loves puzzles. There was a super old, super tricky math puzzle that grown-up mathematicians had been trying to prove for over 350 years! Claude, the robot brain, figured it out in just 11 days. It wrote down the answer in a super long, detailed way, like writing a giant book with 13 million lines, so that even another computer could check every single step to make sure it was right. It's like a robot winning a marathon against the best human runners!
Analysis
The recent achievement by Anthropic's Claude AI in formally proving Fermat's Last Theorem marks a significant leap in artificial intelligence's capacity for abstract reasoning and complex problem-solving. This 358-year-old mathematical enigma, which states that no three positive integers a, b, and c can satisfy the equation aⁿ + bⁿ = cⁿ for any integer value of n greater than 2, was famously proven by Andrew Wiles in 1994. However, Claude's accomplishment lies in generating a fully computer-checkable proof, a feat that took the AI only 11 days and resulted in an unprecedented 13 million lines of code.
Fermat's Last Theorem
Fermat's Last Theorem has long been a benchmark for mathematical rigor and ingenuity. Its initial statement by Pierre de Fermat in 1637, without a published proof, challenged generations of mathematicians. The theorem's eventual proof by Andrew Wiles was a monumental achievement, spanning hundreds of pages and requiring advanced mathematical concepts. Claude's ability to produce a formal, computer-verifiable proof demonstrates a new paradigm in how complex mathematical truths can be established and validated, moving beyond human intuition and potential error.
This development is particularly noteworthy because it provides an irrefutable, line-by-line verification process, a level of certainty that traditional human-generated proofs, while rigorous, cannot always match without extensive peer review. The sheer length of the proof, 13 million lines, highlights the AI's capacity to manage and process an immense amount of logical steps, far exceeding what a human could practically construct or review in a similar timeframe.
Claude AI
Anthropic's Claude AI, the system behind this breakthrough, showcases the power of advanced large language models and AI agents in tackling problems traditionally reserved for human experts. The AI's ability to work largely on its own, as stated by Anthropic, suggests a level of autonomy and understanding that pushes the boundaries of current AI capabilities. This self-sufficiency in generating such a complex and lengthy proof indicates that AI is not merely a tool for computation but is evolving into a partner for discovery and verification in highly specialized fields.
Kevin Buzzard, a mathematician leading a human-led project at Imperial College London aiming to achieve the same formalization, reviewed Claude's proof and confirmed its validity. This endorsement from a human expert in the field lends significant credibility to the AI's output. The fact that Claude beat the human project, which has been running since 2024 and is not close to completion, underscores the efficiency and speed advantages that AI can bring to intellectual endeavors.
Imperial College London
The human-led project at Imperial College London, focused on formally proving Fermat's Last Theorem, serves as a crucial point of comparison for Claude's achievement. This project, led by mathematician Kevin Buzzard, represents the traditional, meticulous approach to formal verification, relying on human intellect and collaborative effort. The ongoing nature of their work since 2024, without reaching completion, highlights the immense time and intellectual resources required for such an undertaking when performed by humans.
The contrast between Claude's 11-day completion and the ongoing human effort at Imperial College London is stark. It illustrates the potential for AI to dramatically accelerate progress in fields like mathematics and computer science, where formal proofs and rigorous verification are paramount. While human oversight and understanding remain essential, AI can act as a powerful accelerator, handling the tedious and complex logical steps that would otherwise consume years of human effort. This collaboration between human expertise and AI efficiency could redefine research methodologies in the future.
Key points
- Anthropic's Claude AI formally proved Fermat's Last Theorem in 11 days.
- The AI generated a 13-million-line proof, making it the longest computer-checkable math proof ever.
- A human-led project at Imperial College London, working on the same task since 2024, is not yet complete.
- Mathematician Kevin Buzzard reviewed and confirmed the validity of Claude's proof.
- This achievement highlights AI's advanced capabilities in complex problem-solving and formal verification.
The ability of AI to generate and verify complex mathematical proofs could significantly enhance the security and reliability of cryptographic algorithms and smart contracts, potentially leading to more robust and trustworthy blockchain systems. This advancement could also accelerate research in areas like zero-knowledge proofs and other advanced cryptographic techniques.
While impressive, the sheer length and complexity of AI-generated proofs, such as the 13 million lines produced by Claude, could make human auditing and comprehension extremely challenging. This might lead to a reliance on AI-generated proofs without full human understanding, potentially introducing subtle, undetected vulnerabilities in critical systems if the AI's logic contains unforeseen flaws.



