Deep Research on a Loop: Agent-Driven Construction of Auditable Structured Datasets
with Santiago Afonso, Sebastian Galiani, and Ramiro H. Gálvez.
Constructing datasets from primary sources remains a bottleneck in empirical research. Deep-research agents can search the web, read documents, and synthesize evidence, but their default output is a narrative memo rather than a comparable, auditable dataset. We introduce Deep Research on a Loop (DRIL), a protocol-driven methodology that turns open-ended agent research into structured data. A design-stage agent helps specify a research instrument, mapped unit space, and evidence policy. These choices are frozen into a protocol that implementation-stage agents apply across units under fixed coding and citation rules. Recorded values must be supported by verbatim evidence, with uncertainty and gaps documented explicitly. We deploy DRIL on a 2025 partial update of a public tax-expenditure database covering 86 analytic jurisdictions, about 40% of the database's universe. The run returns 1,040 sources, 1,113 evidence records, 103 quantitative estimates, and complete qualitative coverage on 22 fields, at about USD 22.50 in equivalent API cost. Quantitative gaps decompose into absent public reporting (71%), document-parsing limits (16%), and residual design-sensitive cases. The results suggest that partial automation can materially reduce the cost of empirical dataset construction when the process is standardized, evidence-bound, and auditable.
PDF
AI Agents and Prompt Engineering in Econometric Coding
with Sebastian Galiani and Federico A. López.
We study how large language models write code for econometric analysis. We compare three dimensions of AI-assisted coding: statistical software (Stata, R, or Python), prompting (zero-shot versus few-shot), and the degree of agency, from a chatbot that writes a single script to an agent that executes and revises its own code. On a benchmark of applied econometric and statistical tasks, moving from the chatbot to the constrained agent raises task success from 74 to 96 percent, at about eight additional cents per run. For Claude Sonnet 4.6 and GPT-5.4 through Codex, few-shot prompting improves the chatbot far more than the constrained agent, indicating that prompting and agency act as substitutes. For these models, differences across statistical software are sizeable under the chatbot but largely disappear under the constrained agent.
PDF
How Institutions Reweight Evidence: Climate Science from the IPCC to the Press
with Sebastian Galiani and Franco Mettola La Giglia.
A climate finding grows more severe as it is summarized for the public, most of it added by the press, not the IPCC. Using three language models, we score this displacement at the claim level across about 1,000 internal and 25,305 media pairs. Both steps lean toward the severe end while staying within accepted ranges. The press step, near +0.17 on a [−1,+1] scale, runs three times the internal one. Left- and right-leaning outlets converge rather than diverge, which fits shared institutional incentives better than partisan demand. The pattern is the systematic outcome of incentivized summarization, not a verdict on the science.
PDF