The Machine That Reruns the Math: Why Cheap Verification Is About to Reprice Trust
Three physicists and an economist built a system that reads economics papers and checks whether the numbers hold. They aimed it at nearly every published paper in five major journals that shipped with a "replication package" — the code and data authors are now required to submit so others can rerun their work. What it found tells you something about economics, and something larger about how AI is about to change every field that runs on unchecked claims.
How the machine actually works
The system is built around a large language model — an AI trained to predict the next chunk of text, which as a side effect can read code, write code, and reason about numbers. Left alone, an LLM would do the dangerous thing: produce a confident, plausible answer with no idea whether it's true. The whole design is built to stop it from guessing.
The mechanism is a loop, and the loop is the point. The AI does not "know" whether a result replicates. It writes a command to run the author's code, the command is actually executed on a real machine, and the output — numbers, or an error message — is fed back to the AI as new text. The AI reads that output the way it reads any prompt and decides the next move: fix a broken file path, install a missing library, compare a number it just produced against the number printed in the paper. It is a text-predictor wired to a computer that gives it real feedback, so its predictions get corrected by reality instead of floating free. That wiring — LLM proposes, machine executes, result returns, LLM reacts — is what separates a useful tool from a hallucination engine, and it is the part worth carrying to every other "can AI do X" question you'll face.
The workflow does three jobs. Keep them separate, because they escalate in difficulty.
Job one: reproduce. Run the published code on the published data, and compare the numbers it produces to the numbers printed in the paper. This sounds trivial. It is not. Code rots — a library updates, a default changes — and a table that once read 0.34 now reruns as 0.31. Did the author round differently, edit the code after making the table, or make a mistake? The system flags every gap between what the code outputs and what the paper claims.
Then it does something the original authors usually didn't: sensitivity analysis — deliberately jostling the result to see if it survives. Drop the extreme data points; change one arbitrary choice the author made, like where to cut off a sample. A robust finding barely moves. A fragile one collapses when you breathe on it. The machine does this automatically, on results no human had the patience to poke.
Job two: improve. Rewrite the author's code to get the identical answer faster — a smarter algorithm, a cleaner implementation. Researchers routinely code the obvious way, not the efficient way, so calculations grind for hours that could take minutes.
Job three: extend. The hardest. Propose a new analysis the paper didn't run — a question consistent with the authors' own data and assumptions but which they never asked. This is nearer to doing research than to checking it.
What it found
Across 4,452 replication packages, the system flagged a discrepancy between the code and the published findings in 3,460 articles — roughly three in four.
Two cautions before you draw a conclusion. First, do not read this as "three-quarters of economics is fraud." Most gaps are small and innocent: rounding, a version mismatch, a figure regenerated after a last code edit. Second — and this is the part the headline number hides — these are flagged by the AI, not adjudicated by a human. The checker is itself a fallible LLM; some of its flags are the machine misreading, not the paper misreporting. You have to hold the tool to the same standard it holds the papers to.
The real lesson survives both caveats: the published number and the number the published code actually produces are routinely not the same number. The replication requirement was supposed to close that gap. It mostly didn't — not from dishonesty, but because nobody had the labor to check thousands of packages line by line. A tireless machine is exactly the labor that was missing.
On speed: in 496 articles, the AI made a calculation run more than ten times faster at equal or better accuracy. When a result takes days of compute, a tenfold cut changes what one researcher can afford to try.
On extension: in 923 articles, it produced a genuine new analysis aligned with the paper's aims — not in the original, not a rehash.
The portable model
Now the idea that outlives economics.
The reliability of any large body of work is set not by how honest its people are, but by how cheap it is to check them.
Peer review makes the point. A few overworked experts read your paper and judge whether the argument sounds right. They almost never rerun your code. The check was on the prose, not the arithmetic — and that was a rational compromise, because rerunning everything was too expensive. The compromise left a standing gap between "reviewed" and "verified" that everyone tolerated, because closing it cost too much.
Watch what happens when the cost of checking falls toward zero. You don't just catch more errors. You change what the field can be trusted to have done. Where checking is cheap — a compiler catching your typo, an airline's maintenance log, a bank reconciliation — standards are high, because errors surface the instant they're made. Where checking is expensive — clinical trial data, government statistics, most academic research — standards quietly drift, not from malice but because the correction mechanism is starved of labor.
The same lever sits under institutions you rely on daily. Legal contracts stuffed with cross-references nobody fully traces. Corporate accounts where auditors sample transactions rather than examine all of them. Engineering specs that "should" match what was actually built. In every case there is a large body of claims and a much smaller amount of real verification, and the gap between them is a cost problem wearing the costume of a trust problem. Cheap automated checking dissolves the cost, and the costume comes off.
So here is what you can now do that you couldn't before reading this. When someone points to a reviewed, audited, or certified body of work and calls it trustworthy, ask the sharper question: was this verified, or merely too expensive to verify? Those have looked identical for as long as checking was costly. They are about to look completely different. The economics result is the early tremor — three in four papers carried a discrepancy not because economists are careless, but because for decades no one could afford to look. Now something can look at all of it, all the time, for almost nothing.
Distilled from NBER Working Papers
Liked this one?
The week's best pieces, one email, every Sunday. Nothing else.