What I Learned Running a Hundred-Page Empirical Study With AI

What I Learned Running a Hundred-Page Empirical Study With AI

Last week I finished a hundred-page empirical report on four federal and two state broadband subsidy programs. It found that the programs mostly paid for service that was coming anyway, and that the ones with credible comparison groups bought a year or so of earlier service at a few thousand dollars per location per year. A project like that would once have taken me a year with a research assistant. It took about two weeks, and no research assistant.

For each program the analysis compared funded areas with similar unfunded ones, built from the programs’ own records where those existed and from unfunded areas with the same starting point where they did not. It tracked both from 2015 to 2025, using data from the FCC, the Census Bureau, two state grant programs, and a private company’s internet speed tests, and it estimated how much of the new service the money caused and what each year of earlier service cost the public.

In May I wrote in the Wall Street Journal about building a retirement planning application with AI in a week. AI had cut out the costly step between knowing what a tool should do and hiring someone to write it. It complemented the expert and substituted for the intermediary. That shift has now reached my own job, and the intermediary it replaced was the research assistant.

Anthropic’s Claude, running as an agent, assembled the data, wrote and ran the code, estimated the regressions, produced the tables and figures, and drafted text. I specified the questions, the designs, and the comparison groups, reviewed the results and the design choices at each step, and read some but not all of the code. OpenAI’s Codex then wrote its own code to re-estimate every family of regressions in the report from the project’s assembled data, and checked substantial parts of the data construction against the written methods. Of 563 numbers in the tables, counting repeats, 561 matched at the precision shown in the tables. The two that did not were rounding errors in converting two coefficients to percentages, and the replication caught them. It did not re-download the raw data or test the causal assumptions. The report says all of this in a note on methods, and I think a note like that should become standard.

My job was directing and checking

I spent the two weeks working like a principal investigator with research assistants who are very fast, very literal, and never tired, but who will go on strike when they need more money to continue working. They do what you ask and nothing more, except when they do something else entirely and report it with confidence, so the job is directing them and then checking. The whole project cost a few hundred dollars in subscription fees and a small overage. At published API prices the same usage would have run to a few thousand dollars for Claude’s work and a few hundred for the Codex replication. Either number is a rounding error next to a research assistant’s salary and the year of my own time the old way would have taken.

I asked for designs, read what came back, and sent it back to them with comments and more questions. The report went through about thirty revisions, each after my own edits to or comments on the one before. Three outside reviews fed in, two from independent agents set up to act as referees and one from the Codex replication. A journal article gets as much review, but of a different kind, and this kind was the only way to know whether the work was right. I expected the main problem to be invented numbers or sources. The checks found none. The errors that reached the referees and the replication were familiar research mistakes hidden in plausible results.

Three mistakes are the most useful part of the story. I caught one of them, late.

The role of independent referee agents

The first mistake was a design error. One analysis dropped every area that later won a newer grant. Those grants went only to places still unserved, so dropping them removed the places where the program had not yet worked. Economists call that selecting on the dependent variable, and a second-year graduate student is trained to avoid it. The agent writing the code made it, and I read the results without noticing. An independent agent acting as a referee caught it, and the fix was simple once flagged. The model makes elementary mistakes along with sophisticated work, with no change in tone to warn you. A second model reviewing the work caught this one cheaply, before a reader did.

The research assistant is literal

The second mistake, which is the one I caught, was mine. Every analysis compared areas that got a subsidy with areas that did not. None used the amount of the subsidy, which left a lot of information unused. My opening prompt asked about the effects of the subsidies, and the model took that literally to mean whether an area got one, not how big it was. I never told it to consider the amount. A human RA would probably have asked, or assumed I wanted dollars. The model does not know what I meant, only what I said, so I have to keep my own theory of mind switched on. Once I asked, the analysis took two hours. In the three programs where the test was possible, larger subsidies per location did not go with larger effects. Grant size was not randomly assigned, so that does not say what an extra dollar would cause, but it is one of the report’s more useful findings, and it exists because asking cost almost nothing.

Looks right versus is right

The third mistake shows the most general advantage of working this way. An early draft of the methods chapter said every program’s coverage shares divided by the same count of locations. In two chapters the code divided by a different one. Neither I nor the referee agents noticed, because the numbers looked right.

In the op-ed I wrote that “looks right” does not mean “is right,” and that closing the gap takes human judgment. In empirical research, judgment is often not enough. A regression table that is a little wrong looks exactly like one that is right, and you cannot reject a result because it surprises you. The usual robustness checks test the design. They do not establish that the code implements the design the text describes.

Catching that means checking what the code does against what the text says it does, and few referees have the time. Codex, in rebuilding the estimates and checking the data construction, did enough of it to find the discrepancy. AI made the mistake, but a human could have made it as easily, and I doubt anyone would have found it. I do not know whether AI is more or less likely than a person to make such a mistake. I do know that AI made a substantial check affordable, and that the check caught what my reading and two referee reviews missed. Applying the stated rule to the same samples moved the estimates by about 0.01 percentage point, so no conclusion changed, and the fix was to the text, not the models. That was luck. The same check would have caught an error that mattered.

Replication is the standard in principle and but in practice rarely happens, because it cost as much as the original work. It now costs a fraction of that. Agreement between two systems is not proof, though. Both can share a flawed input or assumption, so the check needs a stated scope: what was rebuilt, what was reused, and what remains untested. If AI does the empirical work, an independent computational check should be a condition of publishing it, and the disclosure should say who ran it, what it covered, and what it found.

The report’s note on methods says what the model did, what I did, what I did not do, and what the replication checked and found. I remain responsible for the analysis and the interpretation. The disclosure debate in journalism applies with more force to research. In August the Wall Street Journal’s editorial page editor, Paul Gigot, defended running an AI-assisted op-ed by Stanley Druckenmiller without disclosure. The Journal’s own writers may not let AI draft their work, he wrote, but for outside contributors what matters is the argument the author stands behind. A reader of an op-ed wants to know whose views these are. A reader of a regression table needs to know who checked the arithmetic.

Evidence can arrive before the constituency does

Evidence about a program usually arrives late because it is expensive to produce. Programs begun as experiments rarely end, because constituencies form faster than evidence. By the time a subsidy evaluation exists, the recipients, the vendors, and the agency staff have organized around it, and the evidence becomes one more input to a renegotiation rather than a verdict.

Faster analysis does not fix all of that. Outcomes take time to show up in the data, and comparison groups need to be valid. But the same designs work on a program’s first two years of data, so an evaluation of the $42.5 billion BEAD program’s first awards should not wait until 2029. The agencies spending the money face obvious incentive problems when conducting their own evaluations, but publishing award records, rejected applications, and progress reports as the money goes out would let anyone run the evaluation.

My report’s analysis measures deployment, not welfare, and cannot say what a year of earlier service is worth. Its comparison groups are the best the programs’ records allow and are not experiments. A human team would have faced the same limits. AI changed how quickly, and with how few people, I could do the work, and what checks I could afford to require before putting my name on it. If the cost of credible evidence keeps falling, the evidence can arrive before the constituency does.

The report, “Already Being Built: What Rural Broadband Subsidies Paid For, 2018 to 2025,” is available here.

Scott Wallsten is President and Senior Fellow at the Technology Policy Institute, a Policy Fellow at the Stanford Institute for Economic Policy Research (SIEPR), and a senior fellow at the Georgetown Center for Business and Public Policy. He is an economist with expertise in industrial organization and public policy, and his research focuses on competition, regulation, telecommunications, the economics of digitization, and technology policy. He was the economics director for the FCC's National Broadband Plan and has been a lecturer in Stanford University's public policy program, director of communications policy studies and senior fellow at the Progress & Freedom Foundation, a senior fellow at the AEI–Brookings Joint Center for Regulatory Studies, a resident scholar at the American Enterprise Institute, an economist at the World Bank, and a staff economist at the U.S. President's Council of Economic Advisers. He holds a PhD in economics from Stanford University.

Share This Article

artificial intelligence

View More Publications by

Recommended Reads

Already Being Built

Research Roundup – September 2026

TPI Aspen Discussion Recap: AI Finds the Flaws. Can We Fix Them Fast Enough?

Explore More Topics

Antitrust and Competition 188
Artificial Intelligence 48
Broadband 398
Content Moderation 15
Economics and Methods 38
Economics of Digitization 15
Evidence-Based Policy 18
Free Speech 22
Innovation 4
Intellectual Property 56
Miscellaneous 341
Privacy and Security 154
Regulation 26
Trade 2
Uncategorized 5

Related Articles

Already Being Built

Research Roundup – September 2026

TPI Aspen Discussion Recap: AI Finds the Flaws. Can We Fix Them Fast Enough?

The AI Labs Asked for a Waiver. Congress Should Ask for a Justification

Video Now Available: Shane Greenstein on Data Center Economics from TPI Aspen

TPI Aspen Forum 2026: Meeting AI’s Power Demand: Who Pays and Who Builds

Where Does E-Rate Money Actually Go? Now Anyone Can Find Out

AI Productivity and Uptake in Firms, the 2025 AI Agent Index, and More on Labor Market Effects, Research Roundup, July 2026

Sign Up for Updates

This field is for validation purposes and should be left unchanged.

Secret Link