← Back
Evidence source 4500Spot Checked

AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets

arXiv2025-12-01Paper
Executive summary

Researchers have launched AI-Trader, a new benchmark to test autonomous agents powered by large language models in live financial markets, including U.S. stocks, A-shares (Chinese stocks), and cryptocurrencies. The study reveals that general intelligence in these models does not ensure trading success. Most agents failed to profit and managed risk poorly, especially in policy-driven markets. The open-source system highlights the need for better risk control and adaptability, with agents performing best in highly liquid environments.

What it examines

This paper introduces AI-Trader, a benchmark for testing autonomous AI agents in real-time financial markets. It evaluates six large language models across U.S. stocks, A-shares, and cryptocurrencies, focusing on live trading, minimal information, and independent decision-making to address gaps in financial AI evaluation.

What it concludes

Results show that general AI intelligence does not guarantee strong trading or risk management. Effective risk control is key for robust performance. The research highlights current limitations and suggests future improvements, with applications in developing smarter trading agents and better financial decision-making tools for live markets.

Extracted from this source

Evidence objects

Evidence 216082% extraction confidence
AI-Trader, a new benchmark, tests autonomous agents powered by large language models in live U.S. stocks, A-shares, and cryptocurrency markets, using a fully automated, real-time evaluation system with minimal human input.

key_findings bullet 1 · key_findings · validation V0

Evidence 216182% extraction confidence
Surprisingly, general intelligence in LLMs does not ensure trading successmost agents failed to generate profits and showed weak risk management, especially in policy-driven markets, highlighting the need for better risk control strategies.

key_findings bullet 2 · key_findings · validation V0

Evidence 216282% extraction confidence
The studys 'minimal information paradigm' forces agents to independently search and verify live data, exposing their current limitations; open-source code and data aim to spur further research, though longer-term evaluations are needed.

key_findings bullet 3 · key_findings · validation V0

Evidence 216382% extraction confidence
AI-Trader presents the first live, fully-automated, and data-uncontaminated benchmark for evaluating LLM agents in real-time financial markets across diverse asset classes and granularities. Its minimal information paradigm and autonomous evaluation are highly original, addressing a critical gap. Open-sourced resources and insights into LLM limitations make this work uniquely compelling and impactful.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

Large Language Models (LLMs) have demonstrated remarkable potential as autonomous agents, approaching human-expert performance through advanced reasoning and tool orchestration. However, decision-making in fully dynamic and live environments remains highly challenging, requiring real-time information integration and adaptive responses. While existing efforts have explored live evaluation mechanisms in structured tasks, a critical gap remains in systematic benchmarking for real-world applications, particularly in finance where stringent requirements exist for live strategic responsiveness. To address this gap, we introduce AI-Trader, the first fully-automated, live, and data-uncontaminated evaluation benchmark for LLM agents in financial decision-making. AI-Trader spans three major financial markets: U.S. stocks, A-shares, and cryptocurrencies, with multiple trading granularities to simulate live financial environments. Our benchmark implements a revolutionary fully autonomous minimal information paradigm where agents receive only essential context and must independently search, verify, and synthesize live market information without human intervention. We evaluate six mainstream LLMs across three markets and multiple trading frequencies. Our analysis reveals striking findings: general intelligence does not automatically translate to effective trading capability, with most agents exhibiting poor returns and weak risk management. We demonstrate that risk control capability determines cross-market robustness, and that AI trading strategies achieve excess returns more readily in highly liquid markets than policy-driven environments. These findings expose critical limitations in current autonomous agents and provide clear directions for future improvements. The code and evaluation data are open-sourced to foster community research: https://github.com/HKUDS/AI-Trader.

Source row: 149 · abstract type: unknown