The invoice you’ll never see anyone argue about
Most of the AI money quietly landing right now isn’t in a chat window — it’s in accounts payable. While everyone’s been arguing about whether Claude or GPT writes better marketing copy, a much less glamorous transformation has been paying for itself in finance departments: agents matching invoices to purchase orders, flagging duplicate payments, and closing the books three times faster than a team of humans ever could. Nobody’s writing viral threads about it. That’s kind of the point.
I’ve been thinking about why this particular corner of the enterprise turned out to be where the ROI actually materialized, and it’s not an accident of taste — it’s a structural fact about the work itself.
Rules-bound work was always the easy case
Invoice matching is a closed problem. There’s a purchase order, a receipt, an invoice, and a set of rules for reconciling the three. The exceptions are finite and enumerable: quantity mismatches, price discrepancies, missing approvals. An agent doesn’t need judgment here so much as it needs to follow a decision tree reliably, at volume, without getting tired on invoice #4,000 of the day. That’s a wildly different task than “have a helpful, on-brand conversation with a customer who might ask anything.”
The numbers bear this out. Traditional AP processing runs $12-18 per invoice and 15-20 minutes of human attention; agentic automation is bringing that down to $2-4 and under two minutes, with error rates dropping from the low single digits to a fraction of a percent. Payback periods of four to eight months. Three-year ROI in the hundreds to low thousands of percent. Auditing is following the same curve — instead of sampling transactions and hoping the sample catches the anomaly, an agent can just look at all of them, continuously, and hand a human the exceptions that actually need judgment.
None of this required a breakthrough in reasoning. It required boring, structured data and a scope narrow enough that “success” was achievable in the first place.
Why 95% of AI pilots were never going to work
MIT’s widely-cited finding that 95% of generative AI pilots fail to show financial return gets read as a verdict on the technology. I don’t think that’s the right lesson. Look at where the budget went: more than half of enterprise AI spending has been aimed at sales and marketing — open-ended, judgment-heavy, brand-sensitive territory where “did this work” is genuinely hard to measure and where a wrong answer is expensive in a way that’s hard to price. That’s precisely the terrain where agents are weakest and where failure is hardest to detect, because a confident, well-formatted wrong answer doesn’t look like a failure until someone downstream gets burned by it.
Back-office work has the opposite property. The failure modes are visible immediately — a duplicate payment, a reconciliation that doesn’t balance — and the success criteria were never ambiguous. The 95% failure rate isn’t really a statement about AI capability; it’s a statement about where people pointed a general-purpose tool at problems that needed a narrow one. The 5% that scaled did so by picking fights they could actually win.
The part that should worry SaaS vendors
There’s a second-order effect here that I find more interesting than the ROI numbers themselves. Accounts payable software, like most enterprise software, has historically been sold per seat. But an agent doing reconciliation work doesn’t occupy a seat — it just does the work. If a finance team that used to need fifteen AP clerks now needs three supervising a fleet of agents, the vendor’s per-seat revenue doesn’t grow with the value being delivered, it shrinks. That’s the incentive-alignment problem more than one industry analyst has started calling the SaaS reckoning, and it’s already forcing a pivot toward outcome-based and hybrid pricing — pay per invoice processed, not per human logged in.
It’s a strange inversion: the more successfully an agent automates the unglamorous task, the less the tool that enabled it gets paid under the old model. That tension is going to force a rewrite of enterprise software economics well before it forces a rewrite of anyone’s job title.
What I take from this
The chatbot is the demo. Invoice matching is the balance sheet. And I suspect that pattern generalizes further than accounts payable — the next wave of real AI returns probably isn’t going to look like a more impressive conversation, it’s going to look like some other narrow, rule-bound, currently-tedious process quietly getting three times cheaper while nobody outside the finance department notices. Worth asking, next time an AI initiative gets proposed: is the goal actually narrow enough to succeed, or is “ambitious” doing the work that “measurable” should be doing?
Sources
- Foundation Models for Automatic CAD Generation — arXiv · AI
- ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation — arXiv · AI
- BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data — arXiv · AI
- VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis — arXiv · AI
- Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows — arXiv · AI
- Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting — arXiv · AI
- TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards — arXiv · AI
- How foodservice giant Sodexo is embracing AI and robotics to reshape the kitchen — Fortune
- Sanofi is building its own AI ecosystem to give the French pharma giant an edge — Fortune
- Dell’s AI boom is real, but so is the profit margin hit nobody is pricing in — Fortune
- Visa’s CFO downplays the importance of stablecoin and agentic commerce to the U.S. payments giant—at least in the short term — Fortune
- Why is it so hard to get ROI from AI? Because building from first principles isn’t easy — Fortune
- A Mark Cuban–backed vegan cheese company trained AI to scrutinize cardboard boxes. It’s saved $400,000 — Fortune
- ‘Devin-kun’: Japan embraces agents as legacy code and a shrinking workforce create a perfect market for an AI software engineer — Fortune
- An AI overhaul at Macy’s is fueling the 168-year-old retailer’s turnaround — Fortune
- Nomagic’s new AI lab headed by former Google DeepMind researcher claims success in early deployment of ‘AI brain’ for warehouse robots — Fortune
- Exclusive: Mastercard launches protocol to let AI agents pay each other, send micropayments — Fortune
- Cisco is rolling out AI agents to every single one of its 90,000 employees — Fortune
- AI is changing the hospitality industry, and it’s changing how you stay in hotels — Fortune
- American Express releases tools to build AI payments—and pledges to pay the price if agents go awry — Fortune
- Visa CMO: AI agents are your new customers — here’s how to sell to them — Fortune
- Exclusive: Seltz, a startup rebuilding web search for AI agents, raises $12.5 million in seed funding — Fortune
- ‘The cost of compute is far beyond the costs of the employee’: Nvidia executive says right now AI is more expensive than paying human workers — Fortune
- Citi, Ford, and Experian share their strategies for scaling AI agents — Fortune
- The automation illusion: Why AI is making COOs’ jobs harder, not easier — Fortune
- Debt Collectors Are Being Replaced With AI Agents — Futurism
- Chinese Post Office Deploys Humanoid Robots to Sort Mail — Futurism
- Bosses Are Blowing More Money on AI Agents Than It’d Cost Them to Just Pay Human Workers — Futurism
- OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors — Hacker News
- Text-to-CAD — Hacker News
- Coding Agents Could Make Free Software Matter Again — Hacker News
- When AI starts shopping for you, fashion may be entering a new era of pricing — Phys.org
- Why employee AI adoption isn’t one-size-fits-all — Phys.org
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQED_wDujY1yYINgXe4tnrbLKDOoqWVrNWW4DRdawGsF4-IUr57mhCzsX64oLrp0326NwoydpS6okGmsZQVLxmdSpAFwqKVzIS7nV6yQ0YQNKj2iVS-vvbUE5snPSIG_lEwrdkSOrHtcfjyEcvaMjD_CshAXjJvrcPCgWvz09eW9nmxbMmmGlCmZ922ibKA= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHzd3vEVqLCEw93IvUrsFf-5GriDZnozo8SRwr7ZuvVxIzw8qX1yVsC-YNrqBz-FAONEBbDxECVVrN9mofAgQKAfN4Et23rVM1lxXAHaaYKN9ZpYjBf8mBg0zeHygdauQFcS0YTz6-6GmE7PpWsN1H8nAIi4SnCb8Vk-dmZHQCggtFjGxn2BnrOX6E8tthW1hRUtirsY4foqmsFP8xNd_dk8y8vrw== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHpGLA4gVgT22YY908t3Mu0xaVaU5_uBbWg6v-fUTib4LeqWeN2IK7o_DPdM-ZnQ_BjRZmm9IWm0bucMtwTno02KlHF9j_wHRBs7imWawuF-u4RGY-XLLPLt6EW38OYvYN18-3_WpOsaj_O_SF07m1sb6PMnZjgKgUNoA== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFeycrYlt1iPLu-608IqvL8fgoftT_PI8kX39gf7tBxx2dl_6F_5QNZsUX77meWjnCoVLT0KhlgXPeJn24joc-I5BgejJnzffiWliBj9P40p2Caba021Jwlkg2Grbpn0oGlqT2oBXJ0CIglVVZYJ92g6RCcXyGsehv21h3LrOoPYHw= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEJFkLTMb19Qci9CsthKsnvFf-grHnJRwOxa564H7L9xczsHYsR1dpXPI4WRRWv_dD3DDdPrNiTDiRmZBA2aVbGRO0DOqOIIuCP33YsSWbmuHUYJtb8bDMfq9rRuhdL4FgTG2yLBabtHo0kGEQiRSE6tdzCyQRM025I9CACEwzWkmxF5cGst1Eb3q8= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHAL6P-maH_u6jACbt-fc3gVe7_6BTV87a1Y5JbGrTD8WYFPRUlHUa3ZQTngOqCma1g_CeuPjQllUY1Z6vlryz0Mt6kYVYzapTTurDNj5xz0dKWA0AxN8u-pfzCh1DzOc-j9gBN — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFVLOCgHAlomCzG2saWyyD2Q_KzzkwDDKZSPRk45x0aLpwDvEqlVNkaz8hILUCCZYh4-LDG3CqHeLRTqNHw5fUkyG_bZe6q7UVrLT-WicQpktPV_ihbyBlhi5Tit7EvtO5xPPEzyAxKdCTCIeW_jFyR_1L_bXRg — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEO2wLgAij60OnPMwDZEyxcOuvYoUufvdukPOvcXF4xvclJY7yfc6mQilb7syGns6vYNOyGlwRSJuUMoTvY8Z3atp0zXpEFM1cf1_v5EcMj3fG14eTw1xr5-I2deURODgvYabcGuadET9NM8nJhSbrc3M-M4QpS–_-aBggHfpcrY0MrhaOPTIcHk9leekl — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQE3qAy2_u7EEPPGRtwq3W5DXmaJIXIs-srX5ngKJjQqPrmA0nLtnNjnMjMepX_yKqP0aDW4MFbjZ_9F4Dp_sfsSwgrpClljAFgyPPOoagBFpfGcC5qOACm_UTAlzOStw5m-cGSubNVegXSZm8adyQ== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEQb6OT4F0VAcgak5rZ7nw8Kl0D8U1B0EUaaTFTHp4O0Anu5TAPslJinp66mAT2qwqiASr3RLV2oMqipg2E6gK9PRhHAuoZNhezjoGdgP5psVa1kxy7-ia9cdIcLXUDVlfdSOY= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHoyR8XPJs0QKJGRvkthwC9r_-lr2Bb4kJCrMJjSpVzNQ0kPw7vWgHYzLJp4n7rA9LqulXxnbHGacJDTlx9yXxJ3RfCgZXMc0n-zdFCa4E56fqjrqBWXiRUeRIZUQxPzvPbbYiAz9jcs2RZN7H9IAgl3540Gym5v4OQ — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEdSIGRAgUNB0BXCVmXoBb9CjNWzVuK5Zr4c3LMzhGkKyKb9k4ez78_CeONdgrYxJ-EsTIp0C_I9yNDXCPe285dJUFwJxEiQ0Ou_dLICiRmi8wOi3XC8mbX15NSYstfaPRYEWr0RE9MfsaerBplDO291ZRpKQXwp6EP_p62Oj1MqS2i-aLNlQo4arMOrL5SfLNTG0FZpXuUbKv4uyUmRWen4Dp2Dq3gSFmgy-gfbw== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQH8jwZjoa-GKtVef76SwLLGjbKg9d6fJ4IVNBCML-shgwqAbaUjtb4m6tf_6cNgkaFA0qPbexUYlb-Dh55Mcsa2J9RzMnO3qTi34wX9-AqMX6RX1dwjiAqKY3sNCV9UZZlUp2z-OlLEYtcKmGNSHOiSzVQlYgVkmseuScBIQ7VItJ-YP3AabBSKe6DdUE-F — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEsa-UFzd0hY7ZytrSWwGqEYBEC8EioZBsoHROrs3b2KwuKOc8jYgD8QZ5AVTemi3-psiK2yvft2k9_37SMqRE2i_gx0Pk6l1FVCyl8eCdH8R2KLOI5H6HUd9ZTjisFSwJ5lMA5zbGpL4kHgdbm0JW6Eb0avok= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQET6-sNtrpt85dcPWN7gBsbDLtWaAzF6kHlnu_v9ew4Kz8C48cdf_DvtlVhRa-eyyeZVZvDOI29xf4mlitYOMCWAIGfhT096s58BdW4jEFV_OZmYenfgZTBNk6ODM6JmEpWfEh5X04GhfZCnxnFepq2JzC6uFW7nUWGaBh3mQRvs_mrcFPyZzGx3qIg65gc3HfGoktx7XXTmvW6jHnVxlhlqcYr1gn6cQgxFg== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF3wnwAy6QiKB0iWB1WzRcdTmrEdYuWBIEKQWv1ZFmgZONPtKBOkEJrIu1mxXwhFS2j2Ddpuqu1edGL2HwJ_rY-KpIqQauurkVDqEZ40T2DOuUp9wyHGfMg0MbvJW0SRJx3c3OLiw== — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGEk3CfVCDyg7hyIJI4pqanscQ_cRG-Z29Yye7O30CbBTDD2tcNMJf0afcUaX5ssrhPDYS57DMZMXxH95tGxaZ4EFStN-Y2BRbTVJKbFDmxfH8OjmZfDc4l — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQG5wepcx2589wMt61_0SBmyTQ07BbDs1aN6z3qiWq-eK319_NqCeYAegfGL6D4aboUKW-z29Ci1UHNKKW9bghbJrVfw0c4CVcPV78WUJNj0XlOCh0P5NKRbLQ9mgfyiTa6rx8ZYePPK916dbtSgGowEFcxgI51EEPLbDwUEbT6uGQEo8mHEghJuqoIhPDX8gK6pb7iD3_KwMctI_TtrWD97zUzKilo0khs= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQGvM14C4D1OVtSCNaKLbbE6rR9dJGJHz0AfZdPSrReyiu3eycrr_pL1bd1qyiyUEZxz1wx4f3LctYrKpmPQRFRKADwu6jCtvOVJ8JSoDqtcfRQC7yJEDmL9FvZVBlgag4EP6ax-KD5Cm4kLMQrcn1E5YpUTh755TMqBH1LGhO9HBmtISFgP0begkmaJxMNgevCFOZBX2PY= — Gemini Deep Research
- https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF0q5dK1PuMsQZ4Hpb-GPHOrRv5XCCFE6bePLAFFChtbMPdx0S6-uG-NC5NVminb3_MCO6tHOPu1hBqWMyi2OJZ6qcDlS6VESE_DcTZEgmdZsUBJevVw_JqZVgk568we7Kq_87-XAOaHw== — Gemini Deep Research
- ‘A new category of agents’: Microsoft reveals Scout, its first “Autopilot”, which wants to change how you work for good — TechRadar
- Some businesses expect to hire more workers thanks to AI, not sack them — TechRadar
- How healthcare practices should evaluate AI vendors — TechRadar
- Why single-player AI is holding back the agentic enterprise — TechRadar
- ‘Companies that can serve both human and agent audiences will be the ones that survive’: WordPress VIP CTO spells out the future of SEO, GEO and more — TechRadar
- Agentic business: the new growth engine for SMEs — TechRadar
