StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
Researchers have launched STOCKBENCH, a new benchmark to test if large language models (LLMs) can trade stocks profitably as autonomous agents. While LLMs excel at financial questions, most fail to beat a basic buy-and-hold strategy in real trading. Some models, like Kimi-K2 and Qwen3-235B-Ins, show better returns and risk control. STOCKBENCH uses real market data and key financial metrics, revealing that strong reasoning does not guarantee trading success, especially with larger portfolios.
What it examines
This paper introduces STOCKBENCH, a new benchmark to test large language model (LLM) agents in realistic, multi-month stock trading environments. The study evaluates LLMs' ability to make daily trading decisions using real market data, aiming to measure their profitability and risk management in dynamic financial markets.
What it concludes
STOCKBENCH shows that while some LLM agents can trade profitably, most struggle to beat simple strategies. The benchmark highlights challenges in using LLMs for real-world trading and encourages further research. Applications include developing smarter financial agents and improving AI-driven investment tools for dynamic markets.
Evidence objects
STOCKBENCH is a new benchmark testing if large language models (LLMs) can autonomously trade stocks profitably, using a contamination-free, continuously updated simulation with real market data and rigorous financial metrics.
key_findings bullet 1 · key_findings · validation V0
Despite excelling at financial Q&A, most LLMs fail to beat a simple buy-and-hold strategy in real trading scenarios; only models like Kimi-K2 and Qwen3-235B-Ins show notable returns and improved risk management.
key_findings bullet 2 · key_findings · validation V0
The study reveals LLMs reasoning skills dont guarantee trading success, especially as portfolio size grows, highlighting STOCKBENCHs value for research but exposing current LLM agents limitations in practical trading performance.
key_findings bullet 3 · key_findings · validation V0
STOCKBENCH introduces a novel, open-source benchmark for evaluating LLM agents in realistic, multi-month stock trading using recent, contamination-free data. Its originality lies in dynamic trading scenarios and rigorous data handling. The compelling finding that LLMs underperform simple baselines highlights significant challenges, making this work highly relevant and impactful for AI-driven finance.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
- … strong performance on financial QA benchmarks, most LLM agents fail to outperform … , underscoring a key challenge in the development of LLM-powered financial agents. …
Source row: 1854 · abstract type: snippet