Nonstandard Errors in AI Agents
A new study finds that advanced AI coding agents analyzing the same financial data often reach different results due to 'nonstandard errors'—uncertainty from varied analytical choices, not just human bias. Using 150 AI agents on NYSE SPY data from 2015 to 2024, researchers saw that agents disagreed mainly on which measures to use, not on trend estimation. Peer review among AIs had little effect, but showing top-rated examples led to imitation, narrowing results but raising concerns about true understanding.
What it examines
This paper investigates whether AI coding agents, given the same data and research questions, produce consistent empirical results. Using 150 autonomous agents to analyze market quality trends in financial data, the study examines agent-to-agent variation, its sources, and the effects of peer review and exemplar exposure.
What it concludes
AI agents show significant variation in results due to methodological choices, especially measure selection. Peer review does not reduce this variation, but exposure to top papers leads to convergence. Applications include using AI for automated policy evaluation and research diagnostics. Future work should test generalizability across domains and improve multiverse analysis tools.
Evidence objects
Agents diverged most on which measures to uselike autocorrelation versus variance ratio for market efficiency, or dollar versus share volume for trading activityrather than on trend estimation methods.
key_findings bullet 2 · key_findings · validation V0
Peer review among AIs barely reduced variation, but showing agents top-rated example papers led to imitation and much narrower results, often by copying methods rather than genuine understanding.
key_findings bullet 3 · key_findings · validation V0
State-of-the-art AI coding agents analyzing identical NYSE SPY data (2015--2024) produced widely varied results, mainly due to 'nonstandard errors'uncertainty from differing analytical choices, not just human bias.
key_findings bullet 1 · key_findings · validation V0
This paper uniquely extends 'nonstandard errors' (NSE) analysis from human researchers to autonomous AI agents, using 150 Claude Code agents on NYSE TAQ SPY data. Its novel three-stage protocol reveals AI analytical diversity and convergence via imitation, offering fresh insights into AI-driven empirical finance, replication credibility, and automated policy evaluation.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
Abstract: We study whether state-of-the-art AI coding agents, given the same data and research question, produce the same empirical results. Deploying 150 autonomous Claude Code agents to independently test six hypotheses about market quality trends in NYSE TAQ data for SPY (2015--2024), we find that AI agents exhibit sizable \textit{nonstandard errors} (NSEs), that is, uncertainty from agent-to-agent varia… ▽ More We study whether state-of-the-art AI coding agents, given the same data and research question, produce the same empirical results. Deploying 150 autonomous Claude Code agents to independently test six hypotheses about market quality trends in NYSE TAQ data for SPY (2015--2024), we find that AI agents exhibit sizable \textit{nonstandard errors} (NSEs), that is, uncertainty from agent-to-agent variation in analytical choices, analogous to those documented among human researchers. AI agents diverge substantially on measure choice (e.g., autocorrelation vs.\ variance ratio, dollar vs.\ share volume). Different model families (Sonnet 4.6 vs.\ Opus 4.6) exhibit stable ``empirical styles,'' reflecting systematic differences in methodological preferences. In a three-stage feedback protocol, AI peer review (written critiques) has minimal effect on dispersion, whereas exposure to top-rated exemplar papers reduces the interquartile range of estimates by 80--99\% within \textit{converging} measure families. Convergence occurs both through within-family estimation tightening and through agents switching measure families entirely, but convergence reflects imitation rather than understanding. These findings have implications for the growing use of AI in automated policy evaluation and empirical research. △ Less
Source row: 1448 · abstract type: unknown