Finding 7533Emerging EvidenceValidation V0
STOCKBENCH is a new benchmark testing if large language models (LLMs) can autonomously trade stocks profitably, using a contamination-free, continuously updated simulation with real market data and rigorous financial metrics.
86%Confidence
1Evidence objects
v1Version
DraftStatus
Evidence trail
Supporting86% linkage confidence
STOCKBENCH is a new benchmark testing if large language models (LLMs) can autonomously trade stocks profitably, using a contamination-free, continuously updated simulation with real market data and rigorous financial metrics.
key_findings bullet 1 · key_findings
Inspect source: StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? →This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.