← Back
Evidence source 6205Spot Checked

StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?

arXiv2025-10-03Paper
Executive summary

Researchers have launched STOCKBENCH, a new benchmark to test if large language models (LLMs) can trade stocks profitably as autonomous agents. While LLMs excel at financial questions, most fail to beat a basic buy-and-hold strategy in real trading. Some models, like Kimi-K2 and Qwen3-235B-Ins, show better returns and risk control. STOCKBENCH uses real market data and key financial metrics, revealing that strong reasoning does not guarantee trading success, especially with larger portfolios.

What it examines

This paper introduces STOCKBENCH, a new benchmark to test large language model (LLM) agents in realistic, multi-month stock trading environments. The study evaluates LLMs' ability to make daily trading decisions using real market data, aiming to measure their profitability and risk management in dynamic financial markets.

What it concludes

STOCKBENCH shows that while some LLM agents can trade profitably, most struggle to beat simple strategies. The benchmark highlights challenges in using LLMs for real-world trading and encourages further research. Applications include developing smarter financial agents and improving AI-driven investment tools for dynamic markets.

Extracted from this source

Evidence objects

Evidence 753486% extraction confidence
STOCKBENCH is a new benchmark testing if large language models (LLMs) can autonomously trade stocks profitably, using a contamination-free, continuously updated simulation with real market data and rigorous financial metrics.

key_findings bullet 1 · key_findings · validation V0

Evidence 753586% extraction confidence
Despite excelling at financial Q&A, most LLMs fail to beat a simple buy-and-hold strategy in real trading scenarios; only models like Kimi-K2 and Qwen3-235B-Ins show notable returns and improved risk management.

key_findings bullet 2 · key_findings · validation V0

Evidence 753686% extraction confidence
The study reveals LLMs reasoning skills dont guarantee trading success, especially as portfolio size grows, highlighting STOCKBENCHs value for research but exposing current LLM agents limitations in practical trading performance.

key_findings bullet 3 · key_findings · validation V0

Evidence 753786% extraction confidence
STOCKBENCH introduces a novel, open-source benchmark for evaluating LLM agents in realistic, multi-month stock trading using recent, contamination-free data. Its originality lies in dynamic trading scenarios and rigorous data handling. The compelling finding that LLMs underperform simple baselines highlights significant challenges, making this work highly relevant and impactful for AI-driven finance.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

- … strong performance on financial QA benchmarks, most LLM agents fail to outperform … , underscoring a key challenge in the development of LLM-powered financial agents. …

Source row: 1854 · abstract type: snippet