The Proof Is Not the Point Anymore
I’ve started noticing that the code I write matters less than the spec I can’t quite state. I’ll sit down to build something, and the actual coding—the part that used to take hours and require real craft—now takes minutes with an agent doing the typing. What takes the hours now is figuring out, precisely enough that a machine could check it, what I actually meant. The bottleneck moved. It didn’t disappear.
This isn’t just a vibe I have about my own work. It’s the shape of what’s happening at the very top of formal reasoning right now, and mathematics is showing us where the rest of us are headed.
In April 2026, a 23-year-old with no advanced math training fed Erdős Problem #1196—open for almost 60 years—into GPT-5.4 Pro as a single prompt. Eighty minutes later, out came a proof. Not a rehash of known techniques: it introduced something genuinely new, a “downward von Mangoldt Markov chain” that solved the problem by staying inside discrete arithmetic instead of the continuous calculus every human attempt had relied on. Terence Tao and Jared Duker Lichtman spent real effort afterward not proving it, but understanding it—stripping 55 pages of chaotic machine reasoning down into something a person could hold in their head. Lichtman called it a proof from The Book.
That story gets told as a triumph of AI capability, and it is one. But the part I keep turning over is what the humans were doing in it. They weren’t executing logic anymore. They were auditing it, translating a correct-but-alien artifact into something legible, deciding which parts of the insight generalized. Tao’s framing, in a Nature interview earlier this year, is that the mathematician’s job is shifting from hiker to cartographer—not walking the terrain step by step, but surveying what the machine already walked and deciding which parts of the map matter.
Here’s the catch, and it’s the same catch I run into with code: producing a plausible answer got radically cheap, but plausible isn’t correct. The First Proof Challenge—ten fresh, unpublished lemmas handed to AI systems in February 2026, specifically chosen to be impossible to have memorized—showed this starkly. The best public system solved six out of ten. Public consumer models mostly produced confident, fluent, wrong proofs. Confidence and correctness had completely decoupled, and the only way anyone caught the difference was a proof checker like Lean, or a small number of humans willing to grind through the verification by hand. Math has a very rare luxury here: a proof either compiles or it doesn’t. There’s no equivalent oracle for whether the AI-generated invoice-matching script or onboarding flow does what you meant.
That’s the intent gap, and it’s not a new problem—it’s an old problem at a new scale. Code has always been “plausible by construction, not correct by construction.” The difference is that a human reviewing every line used to be the safety net, almost by accident, because someone had to type each one. Once an agent can write, test, and ship a feature end to end, that accidental safety net is gone, and nothing has automatically replaced it. The people working on this—researchers calling it “intent formalization”—are essentially trying to build the Lean checker for ordinary software: postconditions, assertions, formal specs that an AI has to satisfy rather than merely resemble. The honest version of that work admits the real bottleneck isn’t writing the checkable spec, it’s knowing precisely enough what you want to write one. Human intent is messy on purpose. Compressing it into something checkable is the actual skill now, and it’s a skill most of us never had to practice, because we could always just… write the code and see if it felt right.
I think this is why “vibe coding” feels good in the moment and bad a week later. You get the dopamine of a working demo without doing the harder cognitive work of stating, in a form that survives scrutiny, what “working” was supposed to mean. The AI didn’t skip a step you used to do—it skipped the step of you being forced to know what you wanted.
So the question I keep sitting with isn’t “how good will these models get.” It’s simpler and less comfortable: if I can’t state precisely what I want, whose job is it going to be to notice?
Sources
- Prompt-to-Paper: Agentic AI System for Bioinformatics — arXiv · AI
- A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline — arXiv · AI
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science — arXiv · AI
- CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions — arXiv · AI
- LLM Framework for Discovering Major Mathematical Conjectures: AI’s Quest for the Next Riemann Hypothesis — arXiv · AI
- Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics — arXiv · AI
- FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean — arXiv · AI
- Formally Verified Patent Analysis via Dependent Type Theory: Machine-Checkable Certificates from a Hybrid AI + Lean 4 Pipeline — arXiv · AI
- YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate — arXiv · AI
- Effect-Transparent Governance for AI Workflow Architectures: Semantic Preservation, Expressive Minimality, and Decidability Boundaries — arXiv · AI
- Experiments in Agentic AI for Science — arXiv · AI
- Can Generalist Agents Automate Data Curation? — arXiv · AI
- Sound Agentic Science Requires Adversarial Experiments — arXiv · AI
- DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent — arXiv · AI
- OMEGA: Optimizing Machine Learning by Evaluating Generated Algorithms — arXiv · AI
- Nothing from Something: Can a Language Model Discover 0? — arXiv · AI
- Rethinking Publication: A Certification Framework for AI-Enabled Research — arXiv · AI
- Algebraic Semantics of Governed Execution: Monoidal Categories, Effect Algebras, and Coterminous Boundaries — arXiv · AI
- Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes — arXiv · AI
- Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory — arXiv · AI
- Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation — arXiv · AI
- Don’t Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems — arXiv · AI
- LABBench2: An Improved Benchmark for AI Systems Performing Biology Research — arXiv · AI
- Can AI Agents Synthesize Scientific Conclusions? — arXiv · AI
- PrologMCP: A Standardized Prolog Tool Interface for LLM Agents — arXiv · AI
- FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents — arXiv · AI
- First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope — arXiv · Astrophysics
- A vision foundation model for single-cell biology via spatial gene cartography — arXiv · Quantitative Biology
- Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans — arXiv · Quantitative Biology
- TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering — arXiv · Quantitative Biology
- TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology — arXiv · Quantitative Biology
- MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models — arXiv · Quantitative Biology
- Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology — arXiv · Quantitative Biology
- Auditing pretraining contamination in single-cell foundation model benchmarks — arXiv · Quantitative Biology
- Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models — arXiv · Quantitative Biology
- Ontology-constrained multi-LLM scoring of hypothesis support in the predictive processing literature — arXiv · Quantitative Biology
- AI drug discovery leaders warn U.S. health funding cuts risk falling behind global rivals — Fortune
- Zuckerberg Trying to Simulate Human Biology at the Cellular Level — Futurism
- Amateur armed with ChatGPT solves an Erdős problem — Hacker News
- Large language models can predict the results of social science experiments — Nature News
- Towards the construction of a virtual yeast — Nature News
- ‘Virtual cells’ aim to turn raw data into predictive models of biology — Nature News
- AI agents are checking the scientific literature — and spotting decades-old errors — Nature News
- ‘The job description is changing’: mathematician Terence Tao on the rise of AI — Nature News
- A chemistry lab that runs itself to find the perfect reaction — Nature News
- CRISPR gets a power boost from AI-designed ‘molecular scissors’ — Nature News
- How AI is reshaping discovery in maths and physics — Nature News
- AI systems devise hypotheses and ways to test them — Nature News
- Autonomous AI screening flags unreliable Lyme test results, boosting sensitivity to 95.7% — Phys.org
- AI sorts cell droplets into four shapes, uncovering drug effects in human cells — Phys.org
- AI framework could speed battery, combustion and materials research by automating simulations — Phys.org
- Why a tiny social media post has mathematicians rethinking AI — Phys.org
- Finding hidden catalytic knowledge from literature data — Phys.org
- Mathematicians unleash multifold speed boost for supercomputer simulations of molecules — Phys.org
- Physicists and AI model Claude ‘collaborate’ to prove a 10-year-old jamming conjecture — Phys.org
- AI identifies new particle models that may explain neutrinos’ tiny mass — Phys.org
- AI agent helps prepare synchrotron X-ray experimental measurements, paving the way for autonomous operation — Phys.org
- Researchers develop AI tool that finds the equations behind complex systems — Phys.org
- Toward experiment-guided AlphaFold: Researchers overcome AI tool’s single-conformation limitation — Phys.org
- To discover new physics, AI may need to ‘unlearn’ the old one — Phys.org
- Why AI rules in science matter now: Nature backs wider debate beyond mathematics — Phys.org
- AI reveals unexpected source of antibiotic candidates in prion proteins — Phys.org
- AI system translates protein sequences into text, helping reveal functions of unknown proteins — Phys.org
- AI-powered lab discovers brighter lead-free nanomaterials in 12 hours — Phys.org
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHZECweVCpmxNM5AExrH9oE85c0ao2Ev73ncCRRFlIWVG09TdKaanVtspzx1HlKGCW5rHLdHE6ZuihF5SOIgaLK4bsTB8Sg8yt6WurLuFxU_WWgxkq-R9_n-iGCqnWbt6lTk5vOBumofuJmEsPtL8qYS9a75ywmsNWn0ZxXidnCp6sZNvQR98tqqI3vlXg7xnMYiWsa-2HtHAQShJLeMy6OYDiqGAwyfJz8xb0Bk8ryoGEVeUffBZWvVyDw6UpgO8s= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFtkIQtzECA9V_tYWKxpWKsyPkxskHcaYjqrbyc_aB4JssX_j5-TZ5JoqfYjkYhIWquOhnJnq8UJ8OYwF02E1OSOKFxbneXrUidjnkqB-QGpElWBkjRydxDGeL5Falxm3UJl6ldmpIgyy1x7PYXt0ccxnxvlXA_z3gzA67PRVcL8zzgjvzRq4dXfi1AVPBdtKIBtUHAmpOAmCnP7n4vdOQYJ_lPTBjgQg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF3dnHl2NVAgbK_ofsm7JGLXvCAa8QMClwyyWaNuWU-UIQ8wnISOk1qptg7FKbgVigerkY5trwWAZYYWbs2IHo92ETzPErpf8WJbr2F2JupS4doh1-wXdEBJQ== — Gemini Deep Research
- The Uses of Argument in Mathematics — arXiv
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHKlsihardrk6B-pMxuYATy0Qe5XLyWnEEFrFyn4iBGqhqR8tQxaxQJVcefwHn4q9Zja50GgXbvAkboAo2BCFbaD1gXBI2ym7yXBEhbSBkt9co7kyPnJ20Ycfq9YmIl61onhLMtOeYq2d78kPWTzmwq0F7o67I8jdRN_6Mq6_E_Qd1oqAHYzOzVW60Vz89jmQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGCAyYTzEKorHauT3T8ZOPqWvhCt5S5SLiC8ISCq7V0acS98I7Sko_gBtlYbdqy_cn137sQacQITysB7bgBeJwRJzWPcV5C716J5SVaMgdCMjdBxaFWarSzmzamPvgPXuxsgr4bFGOqgnBZ48HsvT6R6xDcB1L415mQ42hyjK1MLrjhQ70dtC0= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEbmj2gU4s7D1ECDv3EnxJ3EhgdUM-AZhdg2HY3Del9fr2cPhNQq2u_DsgFCtbJnqkGViZzVuK50UbqXkVXypFAgCloaTcRfN_2aSf7ganCh9ehUUcZ9ZMIPWzBZGLK_sautf7XhdAzAYD-pENn7ARiuckpb9JBs4OGyj5Wq_CnpvG3OmIzw5f8IuKY — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGlva1PKQwuZnw4zLzD4dKG-xI_7tmn0SmXQTMx2J11JhIx4lDLpbDh8ebeyFRq-tVWQf0qPivelGiFHu1OuwUq1T5p13j6bM-fPylwmcMmXEnxyYkxD0oHSQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFc7aQbOL94e_uoqqV79_WaY-CxaE4RTxQb5tp9ebBU35yFCcIZUXdUhK7OxWl6y9rEkZh1gtFHziYrFUrpUMh5wHiN-scHXilfDkp7x9laSThuOi7Lwvl1BPTLnE33jtKt6Sin2bTj — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEUx_9bkNZRiL4RnXRQrbzeVOppfuCgYicNSd2wDbXFzUtAZq10YLFMXp63B9CfFeBtKcRtA5-lIcDtrAxVzjNQxkf_quJvfG83HIwxzN_sKMhfkk4kCyB22O_Z5G7SJVDddv9ONf7i_0nYVxF07L9921p1gAXxYiyejCH4_CStn9imlpi3 — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEwbI7Wjj0HNywC4dTUYlZjNGAYu-WzqddXQFiigsYtbdQg6bronnkJ8r_XzZfcvSdaBmmF3n1ydc_J4uPW77HO5A1-fz8aGVrGUxfdZJn5HdjnIa876pvY05lcrqZ-qzwYJI7725nBrkPR1oSowDf_VCrMBYYmAR1EdSCNIg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHpEVCdwS9EFaUV2d-IO1OaQGvh0HQg-cwvvp1MVvIcaABk2SMpcNGhMcd3S5hZiT9Y_sgQAQjaAX8o4YR_KNF7LtR9wBHd-PIOgiufDPMCrtgbvOJ2xitxQF8LrEUvhdsnQMDb9JEMFabBfZIJkY8oskTllYhbyaZZk46ZynZl2oOfqPWnJYmFCCSv9CpBNUR9mb3qPlV7RqkWn4sfQI6oaxk= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGjKgfVGMK4rOZgNuxHFd6YhQ6TLYi_KWt_E6rTi1Ws1HkJEPsg8HGtZ3ZtIxG15wJCNZ48qunADqv5ZR9ofYWUNUaGnr6dIafaFVkE1beAs7waC6nwDdaEk6gm_LiREtp_wPZJE0Q4KiRMQ9ZSMGVq — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEuci8mChL9ipnI9iWdP4m_Xsp6tR5ddVZcCC7CCsZctIRaJR4vSnm83VQT9CAH56uOEcEweVr4EymGlVyZ-AfZvUCM6JPFfE_hJe_JjvNMwRhFd7W5wTUc3wVDvVfDfzQecUpicjNzVEF3jX5iU4eVP4wPzWf0NTnx7kMnbXDR64lDo0yWTj-zDBpkLvRYbcOcE3hfE66PKOWm99_j11JOaRqOBQlNzt1RR3Tud3R4ktXhjOQs7jXNomKl0a5JYnPQYfg= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEkiPkH5gGL4gWsZOPMoYBj8keZgViVBUN9IBvawwpz3vLrzAZX67nTN4VygVZzyatMG-pUfKuRTZekk8qrj9dCvu8GXg7Qw33Uu1QpvsTBbWtiMRsLLmHcOw== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQH3-4hnFh89c9PUse6X20YOjyd1cgpNg-9Dbrcl4_kZReyY4CypxWtq5Si9o_9KoMgtWIG7LOQpu_ztuRG-vrfcVQvsLNgmMvGmXg7Zy_fyZo5L0Xc-rjPlli-RuHHwjpCHDynmvsVDJNZHXMA_5qca1_ReF9Mn2SpraAhIy1lDonVOa58jBIA0WY4G9rpW7o8bprk4YloPhyz-Bq8TY4w2P9jvjeSLi7bCv6Y4vzNzJ95hxnyvMRnRKQqRKQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFDAIljNWEfhy0tATADjiIDBBkhHLRGh5cf2Pe49xpKAketPSHtHCmhjAOvKjeGCFE6x7UGEfemNOXI-BRaVVVvrT4tndNB1tl6EnRGVBl2bzm0z54XP5fm37bvFmuIVgupckrssCVZri9TPuPBNKCB_stx6gyr8OhI8Ej4kbRa83tPKNYfTG4cirnm2t0MDyIe-HWUOF4fk6k= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFbThKLHFbxqExnkzgHgF8H8tyYJa8loIUfVU49-_sKrDjgUYnZPSyrLMfGBPF5S4wGTUz1S1gHgXWE5Y7AatqS_iYE-pPtLVGCp8FudVOPMCmwrf15FrfR6k8AUq9-Mdgd — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHeOLDEA0ZHEjlKEjr6As30lL8nRDRKzSnSaVmWEAGOE1xnfkN88RNTu-r55bYMVVoWGYKwnvwDQgikDn-U6ADiZpehDGHc7WeZfHXIAYC_20lN0newKzfOSAb13x89hFutqGfw3qYl5GJISCvh2YZY_6ZzwayGYfskiQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQE6GhJ5iu5Q6wcIrf9KdhFS5Y_lf02wC3PmhFcQ0sL7TO4jK1pKJngVpLo7yadSeE7DRxDQAIQg7bFDoUJEkv6_AMnNuTVi2QLV4F91Z3heJY7e5G2wz6-mbg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHdD5d9HZi6-fiVEsqOcagULcDfi8A2eerVhSkdE8ZQUk7c-VC9fMWCvDI-gUfIPIV_DPwixIW_7OJKnTfnDOKR21bNOr_CxFKngJCwRV3jHfHExfc9uIMMC40CTZxbHp4acCKfUF3xn_0eMmbMXOGOmqX–moPEU74L2dC0dJV42okROCZgIxpD0k69HJdl1x-LztF0oVJUFGxIkZmEbRnKvOEMPvFdswhxdsrZQ== — Gemini Deep Research
- Scientists Turn Mice Transparent to Uncover Obesity’s Secret Effects on Nerves — SciTechDaily
- AI Cracks the Secrets of How the Universe’s Heaviest Elements Are Forged — SciTechDaily
