← Back
Finding 7536Emerging EvidenceValidation V0

STOCKBENCH introduces a novel, open-source benchmark for evaluating LLM agents in realistic, multi-month stock trading using recent, contamination-free data. Its originality lies in dynamic trading scenarios and rigorous data handling. The compelling finding that LLMs underperform simple baselines highlights significant challenges, making this work highly relevant and impactful for AI-driven finance.

86%Confidence
1Evidence objects
v1Version
DraftStatus

Evidence trail

Supporting86% linkage confidence
STOCKBENCH introduces a novel, open-source benchmark for evaluating LLM agents in realistic, multi-month stock trading using recent, contamination-free data. Its originality lies in dynamic trading scenarios and rigorous data handling. The compelling finding that LLMs underperform simple baselines highlights significant challenges, making this work highly relevant and impactful for AI-driven finance.

key_findings bullet 4 · key_findings

Inspect source: StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? →
Knowledge status

This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.