← Back
Evidence source 5197Spot Checked

FinLBench: A Benchmark for Evaluating Large Language Models on Long-Text Financial Documents

Neural Information Processing2025-11-10Paper
Executive summary

Researchers have launched FinLBench, a new benchmark to test how well large language models (LLMs) handle long Chinese financial documents. The core is the FinLEval dataset, featuring 3,219 annotated question-answer pairs from real financial scenarios. FinLBench uses a six-part evaluation system and reveals that commercial LLMs outperform open-source ones. However, all models often invent information, especially with trick questions. This exposes a major flaw in current LLMs and sets a new standard for finance-focused AI evaluation.

What it examines

This paper introduces FinLBench, a benchmark designed to evaluate large language models (LLMs) on understanding and analyzing long Chinese financial documents. It includes a specialized dataset and a six-part evaluation framework, aiming to address the lack of effective tools for assessing LLMs in financial text analysis.

What it concludes

Results show commercial LLMs perform better than open-source ones, but all struggle with hallucinations. FinLBench helps improve LLM evaluation in finance, supporting better model development for tasks like document analysis, risk assessment, and financial decision-making. Future work may focus on reducing hallucinations and expanding benchmark coverage.

Extracted from this source

Evidence objects

Evidence 424778% extraction confidence
FinLBench debuts as a pioneering benchmark for testing large language models ability to analyze long Chinese financial documents, featuring the meticulously annotated FinLEval dataset with 3,219 question-answer pairs across six document types.

key_findings bullet 1 · key_findings · validation V0

Evidence 424878% extraction confidence
Commercial LLMs consistently outperform open-source models on FinLBench tasks, but all models struggle with hallucinationfabricating information, especially when confronted with trap questions, exposing a critical and surprising weakness in current AI systems.

key_findings bullet 2 · key_findings · validation V0

Evidence 424978% extraction confidence
The studys rigorous manual annotation and focus on long-text financial analysis mark a significant advance, though future work should include more diverse documents and address hallucination to further strengthen AI reliability in finance.

key_findings bullet 3 · key_findings · validation V0

Evidence 425078% extraction confidence
FinLBench presents an original, open-source benchmark for evaluating LLMs on long Chinese financial documents, filling a critical gap in financial AI research. Its novelty lies in diverse document types, annotated QA pairs, and empirical analysis of hallucination. This compelling resource advances quantitative finance by enabling standardized, rigorous LLM assessment.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

- The application of large language models (LLMs) in the financial domain is increasing, highlighting the necessity for standardized evaluations. The financial sector …

Source row: 846 · abstract type: snippet