I’ve noticed I spend less time reviewing the code my AI agent writes and more time writing down what “correct” even means before I let it start. That’s a strange sentence to type as someone who spent fifteen years priding himself on writing clean code by hand. But the actual bottleneck in my day has moved somewhere upstream, and I don’t think I built the muscle for where it moved to.
Math got there first
The clearest place to watch this shift happen isn’t in a startup’s codebase — it’s in mathematics, because math has something almost no other field has: a way to check an answer without asking a human. A proof written in Lean either compiles or it doesn’t. There’s no code review, no staging environment, no “looks right to me.” That property is what makes math the canary for what’s coming everywhere else.
In April 2026, a 23-year-old with no formal math training fed Erdős Problem #1196 — open for almost sixty years — into GPT-5.4 Pro as a single prompt. Eighty minutes later, out came a proof. Not a rehash of known techniques either: it introduced a genuinely new object, a “downward von Mangoldt Markov chain,” that sidestepped the precision losses that had stumped human mathematicians working in continuous calculus. The model stayed in discrete arithmetic the whole time and found a mechanism nobody had written down before.
Here’s the part that matters more than the proof itself, though. The raw output was 55 pages of chaotic, barely-parseable reasoning. It took Terence Tao and Jared Duker Lichtman — genuine Fields-Medal-caliber minds — to go in and distill it into something a human could actually hold in their head. Tao didn’t write the proof. He edited it. He decided which parts of the machine’s reasoning were the real insight and which were noise. That’s not calculation. That’s curation, and it’s a completely different skill than the one graduate students spend a decade building.
Plausible isn’t the same as correct
This is where it connects back to code, and where it gets less comfortable. Researchers at Microsoft describe something they call the “intent gap” — the distance between what you meant and what the AI actually built. The uncomfortable line from their work: AI-generated code is “plausible by construction, not correct by construction.” It looks right. It compiles. It passes the obvious tests. Whether it does what you actually wanted is a separate question entirely, and increasingly nobody is answering it, because the whole point of an agentic tool is that you stopped reading every line.
In the old world, that was fine, because a human reviewed each diff and the review was the specification check. Take the human out of that loop — which is the entire premise of the agent products we’re all shipping right now — and the safeguard just disappears. Nothing replaces it unless you build something to replace it.
The replacement, it turns out, is writing down intent as something checkable. That’s a spectrum, not a switch: on the light end, a test suite that pins down the ambiguous edges of a vague prompt. On the heavy end, actual formal specifications in verification-aware languages — pre- and post-conditions the compiler enforces, not just runs. I’ve started doing a version of the light end reflexively before I hand anything nontrivial to an agent, and it’s slower going in and dramatically faster coming out, because the failure mode shifts from “silently wrong” to “won’t compile.”
Grounding, not hype
It would be easy to read the Erdős story and conclude autonomous AI research is basically solved. It isn’t. A month after that proof, eleven mathematicians ran an experiment called the First Proof Challenge: ten genuinely unpublished lemmas, guaranteed absent from any training set, one week, no help. Public models solved two, and mostly the two that resembled things they’d seen before. The best system, from a research group at ETH Zurich, got six. Private frontier models claimed six too, with real disagreement about whether “autonomous” was even the right word for how they got there.
So the capability is real and the ceiling is much lower than the headline result suggests. Both things are true. What’s consistent across both the triumph and the near-misses is where the human work actually landed — not in the step-by-step execution, but in deciding what counted as a real result, auditing the output, and knowing which direction was worth pointing the machine in at all.
I keep coming back to the word cartographer, which Tao used to describe the new job description. Not the one walking the terrain — the one deciding which terrain is worth mapping, and trusting the drone to do the walking. I didn’t train for that job. Almost none of us did. The question I don’t have a clean answer to yet is whether “formalizing intent” is a skill you can pick up the way you picked up a new language, or whether it’s closer to mathematical taste — something that only shows up after you’ve done the slow, unassisted version long enough to know what you’re actually asking for.
Sources
- Prompt-to-Paper: Agentic AI System for Bioinformatics — arXiv · AI
- A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline — arXiv · AI
- Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science — arXiv · AI
- CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions — arXiv · AI
- Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics — arXiv · AI
- FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean — arXiv · AI
- Formally Verified Patent Analysis via Dependent Type Theory: Machine-Checkable Certificates from a Hybrid AI + Lean 4 Pipeline — arXiv · AI
- YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate — arXiv · AI
- Experiments in Agentic AI for Science — arXiv · AI
- Can Generalist Agents Automate Data Curation? — arXiv · AI
- Sound Agentic Science Requires Adversarial Experiments — arXiv · AI
- DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent — arXiv · AI
- OMEGA: Optimizing Machine Learning by Evaluating Generated Algorithms — arXiv · AI
- Nothing from Something: Can a Language Model Discover 0? — arXiv · AI
- Rethinking Publication: A Certification Framework for AI-Enabled Research — arXiv · AI
- Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes — arXiv · AI
- Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory — arXiv · AI
- Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation — arXiv · AI
- Don’t Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems — arXiv · AI
- LABBench2: An Improved Benchmark for AI Systems Performing Biology Research — arXiv · AI
- Can AI Agents Synthesize Scientific Conclusions? — arXiv · AI
- PrologMCP: A Standardized Prolog Tool Interface for LLM Agents — arXiv · AI
- FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents — arXiv · AI
- First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope — arXiv · Astrophysics
- A vision foundation model for single-cell biology via spatial gene cartography — arXiv · Quantitative Biology
- Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans — arXiv · Quantitative Biology
- TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering — arXiv · Quantitative Biology
- TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology — arXiv · Quantitative Biology
- MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models — arXiv · Quantitative Biology
- Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology — arXiv · Quantitative Biology
- Auditing pretraining contamination in single-cell foundation model benchmarks — arXiv · Quantitative Biology
- Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models — arXiv · Quantitative Biology
- Ontology-constrained multi-LLM scoring of hypothesis support in the predictive processing literature — arXiv · Quantitative Biology
- AI drug discovery leaders warn U.S. health funding cuts risk falling behind global rivals — Fortune
- Zuckerberg Trying to Simulate Human Biology at the Cellular Level — Futurism
- Amateur armed with ChatGPT solves an Erdős problem — Hacker News
- Large language models can predict the results of social science experiments — Nature News
- Towards the construction of a virtual yeast — Nature News
- ‘Virtual cells’ aim to turn raw data into predictive models of biology — Nature News
- ‘The job description is changing’: mathematician Terence Tao on the rise of AI — Nature News
- A chemistry lab that runs itself to find the perfect reaction — Nature News
- CRISPR gets a power boost from AI-designed ‘molecular scissors’ — Nature News
- How AI is reshaping discovery in maths and physics — Nature News
- AI systems devise hypotheses and ways to test them — Nature News
- Autonomous AI screening flags unreliable Lyme test results, boosting sensitivity to 95.7% — Phys.org
- AI sorts cell droplets into four shapes, uncovering drug effects in human cells — Phys.org
- AI framework could speed battery, combustion and materials research by automating simulations — Phys.org
- Why a tiny social media post has mathematicians rethinking AI — Phys.org
- Finding hidden catalytic knowledge from literature data — Phys.org
- Mathematicians unleash multifold speed boost for supercomputer simulations of molecules — Phys.org
- Physicists and AI model Claude ‘collaborate’ to prove a 10-year-old jamming conjecture — Phys.org
- AI identifies new particle models that may explain neutrinos’ tiny mass — Phys.org
- AI agent helps prepare synchrotron X-ray experimental measurements, paving the way for autonomous operation — Phys.org
- Researchers develop AI tool that finds the equations behind complex systems — Phys.org
- Toward experiment-guided AlphaFold: Researchers overcome AI tool’s single-conformation limitation — Phys.org
- To discover new physics, AI may need to ‘unlearn’ the old one — Phys.org
- Why AI rules in science matter now: Nature backs wider debate beyond mathematics — Phys.org
- AI reveals unexpected source of antibiotic candidates in prion proteins — Phys.org
- AI system translates protein sequences into text, helping reveal functions of unknown proteins — Phys.org
- AI-powered lab discovers brighter lead-free nanomaterials in 12 hours — Phys.org
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHZECweVCpmxNM5AExrH9oE85c0ao2Ev73ncCRRFlIWVG09TdKaanVtspzx1HlKGCW5rHLdHE6ZuihF5SOIgaLK4bsTB8Sg8yt6WurLuFxU_WWgxkq-R9_n-iGCqnWbt6lTk5vOBumofuJmEsPtL8qYS9a75ywmsNWn0ZxXidnCp6sZNvQR98tqqI3vlXg7xnMYiWsa-2HtHAQShJLeMy6OYDiqGAwyfJz8xb0Bk8ryoGEVeUffBZWvVyDw6UpgO8s= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFtkIQtzECA9V_tYWKxpWKsyPkxskHcaYjqrbyc_aB4JssX_j5-TZ5JoqfYjkYhIWquOhnJnq8UJ8OYwF02E1OSOKFxbneXrUidjnkqB-QGpElWBkjRydxDGeL5Falxm3UJl6ldmpIgyy1x7PYXt0ccxnxvlXA_z3gzA67PRVcL8zzgjvzRq4dXfi1AVPBdtKIBtUHAmpOAmCnP7n4vdOQYJ_lPTBjgQg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF3dnHl2NVAgbK_ofsm7JGLXvCAa8QMClwyyWaNuWU-UIQ8wnISOk1qptg7FKbgVigerkY5trwWAZYYWbs2IHo92ETzPErpf8WJbr2F2JupS4doh1-wXdEBJQ== — Gemini Deep Research
- The Uses of Argument in Mathematics — arXiv
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHKlsihardrk6B-pMxuYATy0Qe5XLyWnEEFrFyn4iBGqhqR8tQxaxQJVcefwHn4q9Zja50GgXbvAkboAo2BCFbaD1gXBI2ym7yXBEhbSBkt9co7kyPnJ20Ycfq9YmIl61onhLMtOeYq2d78kPWTzmwq0F7o67I8jdRN_6Mq6_E_Qd1oqAHYzOzVW60Vz89jmQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGCAyYTzEKorHauT3T8ZOPqWvhCt5S5SLiC8ISCq7V0acS98I7Sko_gBtlYbdqy_cn137sQacQITysB7bgBeJwRJzWPcV5C716J5SVaMgdCMjdBxaFWarSzmzamPvgPXuxsgr4bFGOqgnBZ48HsvT6R6xDcB1L415mQ42hyjK1MLrjhQ70dtC0= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEbmj2gU4s7D1ECDv3EnxJ3EhgdUM-AZhdg2HY3Del9fr2cPhNQq2u_DsgFCtbJnqkGViZzVuK50UbqXkVXypFAgCloaTcRfN_2aSf7ganCh9ehUUcZ9ZMIPWzBZGLK_sautf7XhdAzAYD-pENn7ARiuckpb9JBs4OGyj5Wq_CnpvG3OmIzw5f8IuKY — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGlva1PKQwuZnw4zLzD4dKG-xI_7tmn0SmXQTMx2J11JhIx4lDLpbDh8ebeyFRq-tVWQf0qPivelGiFHu1OuwUq1T5p13j6bM-fPylwmcMmXEnxyYkxD0oHSQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFc7aQbOL94e_uoqqV79_WaY-CxaE4RTxQb5tp9ebBU35yFCcIZUXdUhK7OxWl6y9rEkZh1gtFHziYrFUrpUMh5wHiN-scHXilfDkp7x9laSThuOi7Lwvl1BPTLnE33jtKt6Sin2bTj — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEUx_9bkNZRiL4RnXRQrbzeVOppfuCgYicNSd2wDbXFzUtAZq10YLFMXp63B9CfFeBtKcRtA5-lIcDtrAxVzjNQxkf_quJvfG83HIwxzN_sKMhfkk4kCyB22O_Z5G7SJVDddv9ONf7i_0nYVxF07L9921p1gAXxYiyejCH4_CStn9imlpi3 — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEwbI7Wjj0HNywC4dTUYlZjNGAYu-WzqddXQFiigsYtbdQg6bronnkJ8r_XzZfcvSdaBmmF3n1ydc_J4uPW77HO5A1-fz8aGVrGUxfdZJn5HdjnIa876pvY05lcrqZ-qzwYJI7725nBrkPR1oSowDf_VCrMBYYmAR1EdSCNIg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHpEVCdwS9EFaUV2d-IO1OaQGvh0HQg-cwvvp1MVvIcaABk2SMpcNGhMcd3S5hZiT9Y_sgQAQjaAX8o4YR_KNF7LtR9wBHd-PIOgiufDPMCrtgbvOJ2xitxQF8LrEUvhdsnQMDb9JEMFabBfZIJkY8oskTllYhbyaZZk46ZynZl2oOfqPWnJYmFCCSv9CpBNUR9mb3qPlV7RqkWn4sfQI6oaxk= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGjKgfVGMK4rOZgNuxHFd6YhQ6TLYi_KWt_E6rTi1Ws1HkJEPsg8HGtZ3ZtIxG15wJCNZ48qunADqv5ZR9ofYWUNUaGnr6dIafaFVkE1beAs7waC6nwDdaEk6gm_LiREtp_wPZJE0Q4KiRMQ9ZSMGVq — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEuci8mChL9ipnI9iWdP4m_Xsp6tR5ddVZcCC7CCsZctIRaJR4vSnm83VQT9CAH56uOEcEweVr4EymGlVyZ-AfZvUCM6JPFfE_hJe_JjvNMwRhFd7W5wTUc3wVDvVfDfzQecUpicjNzVEF3jX5iU4eVP4wPzWf0NTnx7kMnbXDR64lDo0yWTj-zDBpkLvRYbcOcE3hfE66PKOWm99_j11JOaRqOBQlNzt1RR3Tud3R4ktXhjOQs7jXNomKl0a5JYnPQYfg= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEkiPkH5gGL4gWsZOPMoYBj8keZgViVBUN9IBvawwpz3vLrzAZX67nTN4VygVZzyatMG-pUfKuRTZekk8qrj9dCvu8GXg7Qw33Uu1QpvsTBbWtiMRsLLmHcOw== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQH3-4hnFh89c9PUse6X20YOjyd1cgpNg-9Dbrcl4_kZReyY4CypxWtq5Si9o_9KoMgtWIG7LOQpu_ztuRG-vrfcVQvsLNgmMvGmXg7Zy_fyZo5L0Xc-rjPlli-RuHHwjpCHDynmvsVDJNZHXMA_5qca1_ReF9Mn2SpraAhIy1lDonVOa58jBIA0WY4G9rpW7o8bprk4YloPhyz-Bq8TY4w2P9jvjeSLi7bCv6Y4vzNzJ95hxnyvMRnRKQqRKQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFDAIljNWEfhy0tATADjiIDBBkhHLRGh5cf2Pe49xpKAketPSHtHCmhjAOvKjeGCFE6x7UGEfemNOXI-BRaVVVvrT4tndNB1tl6EnRGVBl2bzm0z54XP5fm37bvFmuIVgupckrssCVZri9TPuPBNKCB_stx6gyr8OhI8Ej4kbRa83tPKNYfTG4cirnm2t0MDyIe-HWUOF4fk6k= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFbThKLHFbxqExnkzgHgF8H8tyYJa8loIUfVU49-_sKrDjgUYnZPSyrLMfGBPF5S4wGTUz1S1gHgXWE5Y7AatqS_iYE-pPtLVGCp8FudVOPMCmwrf15FrfR6k8AUq9-Mdgd — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHeOLDEA0ZHEjlKEjr6As30lL8nRDRKzSnSaVmWEAGOE1xnfkN88RNTu-r55bYMVVoWGYKwnvwDQgikDn-U6ADiZpehDGHc7WeZfHXIAYC_20lN0newKzfOSAb13x89hFutqGfw3qYl5GJISCvh2YZY_6ZzwayGYfskiQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQE6GhJ5iu5Q6wcIrf9KdhFS5Y_lf02wC3PmhFcQ0sL7TO4jK1pKJngVpLo7yadSeE7DRxDQAIQg7bFDoUJEkv6_AMnNuTVi2QLV4F91Z3heJY7e5G2wz6-mbg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHdD5d9HZi6-fiVEsqOcagULcDfi8A2eerVhSkdE8ZQUk7c-VC9fMWCvDI-gUfIPIV_DPwixIW_7OJKnTfnDOKR21bNOr_CxFKngJCwRV3jHfHExfc9uIMMC40CTZxbHp4acCKfUF3xn_0eMmbMXOGOmqX–moPEU74L2dC0dJV42okROCZgIxpD0k69HJdl1x-LztF0oVJUFGxIkZmEbRnKvOEMPvFdswhxdsrZQ== — Gemini Deep Research
- AI Cracks the Secrets of How the Universe’s Heaviest Elements Are Forged — SciTechDaily
