Forget the Argument, Check the Output

In September 2026, Anthropic announced that its Claude model had turned a famous, decades old mathematical proof into code a computer can check line by line. The proof was Andrew Wiles' 1995 solution to Fermat's Last Theorem, one of the most celebrated results in modern mathematics. Claude converted it into 13 million lines of Lean, a programming language built for formal verification, in eleven days. It did this by spinning up dozens of AI agents that chewed through 29,500 smaller theorems along the way [SiliconAngle]. Mathematicians had expected this kind of formalization to take years of very human, very caffeinated labor.

That's one data point. It doesn't prove AI is smart, useful, or safe, and it definitely doesn't settle the "is this a bubble" debate happening in every group chat right now. But here's the nice part: it's checkable. Anyone with the training can open the Lean file and confirm the code compiles and the proof holds, no faith required. That puts it in a different category from most AI claims you'll read today, which usually come from someone with skin in the game: a company selling the technology, a rival hoping it flops, or a critic who already wrote the "I told you so" post. This article sticks to what AI has actually produced in three fields of science, using only results other experts have checked, tested, or poked at with their own equipment. What you make of it is entirely up to you, not this article.

Why Science Is a Fair Test

Most arguments about AI happen in places where nobody has to prove anything: opinion columns, group chats, investor calls. Science plays by different rules. A mathematical proof compiles or it doesn't. A drug candidate kills bacteria in a lab dish or it doesn't. A cosmological model matches new telescope data or it doesn't. That makes science one of the rare arenas where an AI's output gets graded by people who don't care whether AI wins, only whether the specific claim is true.

The same pattern shows up in every example below: an AI proposes something, a proof, a molecule, a measurement, and human experts check it before anyone treats it as real. This isn't AI going off and discovering things alone in a locked room. It's closer to a very fast, very tireless intern who still needs a manager to sign off before anything ships.

Mathematics: Proof Checking and Problem Solving at Machine Speed

Formal mathematics happens to be a great match for AI, because a proof written in a language like Lean gets checked automatically by a computer. No arguing, no "well, actually," just a compile that passes or fails. Claude's formalization of Fermat's Last Theorem is the clearest recent example. It didn't discover a new proof; it translated Wiles' existing one into a form a machine could verify, leaning on mathematician Kevin Buzzard's earlier groundwork and an open source planning tool called Prove2Me [SiliconAngle]. The result is now the largest Lean file ever written, which is either a proud achievement or a genuinely unreasonable file size, depending on your relationship with your code editor.

A separate, weirder result came from OpenAI, whose general purpose language model, not a system built specifically for math, disproved a 1946 conjecture by Paul Erdős about points on a flat surface spaced exactly one unit apart. The model linked the geometry problem to an unrelated field, algebraic number theory, and used a mathematical tool from 1964 to build a counterexample that simple grid arrangements never could reach. Nine mathematicians, including Tim Gowers, a Fields Medal winner (mathematics' version of an Olympic gold), checked and refined the work before it was published [Memeburn]. Disproving a conjecture, finding one case that breaks it, is generally easier than proving a new result true in every case, so call this impressive rather than earth shattering.

Both cases sit inside a wider trend. In 2024, Google DeepMind's AlphaProof and AlphaGeometry solved International Mathematical Olympiad problems at a silver medal standard, missing gold by a single point out of 609 competitors [DeepMind]. The systems needed each problem manually translated into formal language first, and took far longer than the four hours human teenagers get. Still, three separate labs producing checkable math results within about two years suggests an actual, repeatable skill, not a one off party trick.

pexels-silverkblack-39190909.jpg

Biology: From Predicting Structures to Proposing Drugs

Biology's clearest AI win predates this year by a while, but it's still the strongest card in the deck. In 2024, Demis Hassabis and John Jumper won half of the Nobel Prize in Chemistry for AlphaFold, an AI system that predicts a protein's three dimensional shape from its genetic sequence, a problem biologists had chipped away at for fifty years. (David Baker took the other half for solving the reverse problem: designing brand new proteins from scratch.) The AlphaFold database has grown from about 360,000 predicted structures in 2021 to over 200 million today, covering more than a million organisms, and over a million researchers have used it to study diseases including COVID 19 [EMBL].

A more recent case takes the same idea and points it at drug discovery. In April 2026, researchers at McMaster University, working with Stanford collaborators, used a generative AI model called SyntheMol RL to search through 46 billion possible chemical compounds, more than any human lab could ever physically screen, hunting for new antibiotics against drug resistant staph bacteria. The model proposed 79 candidates, and one, named synthecin, worked as a topical treatment against resistant staph infections in mice [Technology Networks]. The researchers are upfront that this is still an early step: nobody yet knows how synthecin actually works, it's only been tested on mice, and an approved medicine is still years away, so don't go looking for it at the pharmacy just yet.

Cosmology: Squeezing More Signal Out of Billion Dollar Data

pace telescopes and galaxy surveys cost a fortune to build, so squeezing more information out of data that's already been collected has real value. A team at Princeton University and the Flatiron Institute built a method called SimBIG, which trains AI models on millions of simulated universes and then turns them loose on real galaxy survey data. SimBIG found small scale galaxy clustering patterns that older statistical methods simply couldn't see, cutting the uncertainty in key cosmic measurements by more than half, roughly the same gain you'd get from analyzing four times as much data as was actually collected [The Brighter Side of News]. Given that a single galaxy survey can cost hundreds of millions of dollars, that's not a small trick.

A second team, from the Harvard Smithsonian Center for Astrophysics and the University of Tokyo, pointed a similar approach at dark energy, the mysterious force apparently speeding up the universe's expansion. They trained a neural network on light distortions caused by "cannibal stars," stars that have snacked on a companion star, as seen by the Euclid Space Telescope, and used it to measure dark energy's effects about 10 percent more precisely than before [Archyde]. Neither team is claiming to have solved dark energy or ended cosmology's bigger arguments; both describe their AI methods as better tools for squeezing insight out of observations, with results that still need confirming from upcoming surveys.

pexels-tomas-anunziata-129267-695477 (1).jpg

What These Cases Have in Common, and What They Don't Prove

Line up all six examples and the same shape keeps showing up. An AI proposes a proof, a molecule, or a measurement. A group of human experts, mathematicians, lab researchers, astronomers, then checks it before anyone treats it as real. None of these results happened because an AI worked alone in the dark; they happened because an AI did a mountain of preliminary legwork much faster than humans could, then handed the results to people qualified to say yes or no.

None of it is as tidy as the headlines make it sound, either. Formalizing a 1995 proof isn't the same as discovering new mathematics. Disproving a conjecture isn't the same as proving a hard new result. A compound that clears an infection in a mouse is not an approved drug, however hopeful the mouse looks. A 10 percent precision gain in one dark energy measurement doesn't settle cosmology's open arguments. Taken one at a time, each story is real but modest. Taken together, they show the same capability turning up across totally unrelated fields within about two years, and that part is worth paying attention to.

Decide From the Evidence, Not From Whoever Is Talking

Ask ten people whether AI is a bubble and you'll get eleven opinions, most delivered with total confidence. That's because almost everyone talking about AI has a horse in the race: the company selling it wants you amazed, the rival betting against it wants you embarrassed for being amazed, the pundit who called "bubble" years ago just wants to be right eventually. None of them are lying, exactly. They're just invested.

You're not being paid by any of them. Lucky you. So skip the shouting match and look at what nobody can argue with: a Lean file that compiles or doesn't, a mouse that gets better or doesn't, a measurement that holds up or doesn't. Six examples, six receipts, zero opinions required.

Be the annoying friend who checks the source at dinner table AI debates. Read what these systems actually did, try the tools yourself, and land wherever you land, bubble, breakthrough, or something boringly in between. It's your call, not anyone else's talking points.

References

Anthropic Uses Claude to Formalize Proof of Fermat's Last Theorem, SiliconAngle

AI Math Breakthrough: OpenAI Disproves 80-Year Erdős Conjecture, Memeburn

AI achieves silver-medal standard solving International Mathematical Olympiad problems, Google DeepMind

Computational protein design and protein structure prediction win Nobel Prize in Chemistry, EMBL

From Algorithms to Antimicrobials: AI Model Generates Novel Antibiotic Compound, Technology Networks

AI breakthrough unlocks hidden patterns in the universe's structure, The Brighter Side of News

Scientists Use AI & Cannibal Stars to Unlock Dark Energy's Cosmic Mystery, Archyde