← Back
Evidence source 4667Spot Checked

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

arXiv2026-06-02Paper
Executive summary

Researchers have launched BIGFINANCEBENCH, a new benchmark to test AI agents on real-world financial research tasks, not just final answers. The study reveals that top AI models score below 60 percent on detailed workflows and under 45 percent on final answers, far from professional standards. The benchmark features 928 expert-written questions with step-by-step rubrics, exposing that most AI errors occur during information gathering and setup. It currently covers only public-company, English-language tasks.

What it examines

This paper introduces BIGFINANCEBENCH, a benchmark of 928 expert-written financial research questions. It evaluates financial-research agents on realistic, auditable workflows using point-weighted rubrics, focusing on the full process of deriving answers, not just final results. The goal is to measure agent performance in real-world finance tasks.

What it concludes

Results show current AI agents are far from mastering full financial workflows, with best models scoring below 60%. BIGFINANCEBENCH helps identify strengths and weaknesses, supporting safer, more transparent financial automation. Applications include model evaluation, training, and routing. Limitations include static, public-source data; future work may expand coverage and improve agent selection.

Extracted from this source

Evidence objects

Evidence 268186% extraction confidence
BIGFINANCEBENCH debuts as a pioneering benchmark, evaluating financial-research AI agents on real-world, auditable workflows with 928 expert-authored questions and a point-weighted rubric, shifting focus from simple Q&A to detailed financial reasoning.

key_findings bullet 1 · key_findings · validation V0

Evidence 268286% extraction confidence
Surprisingly, even top AI models scored below 60% on the detailed rubric and under 45% on final-answer accuracy, exposing a significant gap between current AI capabilities and professional financial analysis standards.

key_findings bullet 2 · key_findings · validation V0

Evidence 268386% extraction confidence
The study reveals most AI errors occur during information retrieval and setup, not calculation, and highlights limitations: tasks are restricted to public-company, English-language scenarios, excluding proprietary databases and live analyst interactions.

key_findings bullet 3 · key_findings · validation V0

Evidence 268486% extraction confidence
BIGFINANCEBENCH is an original, workflow-grounded benchmark evaluating financial-research agents by auditing derivation processes, not just final answers. Its novelty lies in open-ended, multi-source, assumption-dependent tasks and point-weighted rubrics. This compelling approach reveals significant model limitations, driving impactful research in AI, LLMs, and Machine Learning for Investment Management and Trading.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were used, which assumptions were made, and how the calculation was performed. Existing finance benchmarks largely evaluate isolated subskills or final answers, leaving the auditable derivation itself under-measured. We introduce BigFinanceBench, a 928-item expert-authored benchmark of open-ended financial-research tasks in which each item pairs a ground-truth reference answer with a point-weighted rubric that decomposes the derivation into independently checkable steps. BigFinanceBench is workflow-grounded in that it evaluates the full derivation rather than only the final output. Across 36,241 rubric points, the benchmark supports partial-credit evaluation and localization of failures across the analyst workflow. Evaluating ten current frontier and open-weight agents, we find substantial headroom: the best system reaches only 58.8% rubric score, final-answer accuracy is a useful but lossy proxy for derivation quality, and model capability varies non-uniformly across financial workflows.

Source row: 316 · abstract type: unknown