Finding 4246Emerging EvidenceValidation V0
FinLBench debuts as a pioneering benchmark for testing large language models ability to analyze long Chinese financial documents, featuring the meticulously annotated FinLEval dataset with 3,219 question-answer pairs across six document types.
78%Confidence
1Evidence objects
v1Version
DraftStatus
Evidence trail
Supporting78% linkage confidence
FinLBench debuts as a pioneering benchmark for testing large language models ability to analyze long Chinese financial documents, featuring the meticulously annotated FinLEval dataset with 3,219 question-answer pairs across six document types.
key_findings bullet 1 · key_findings
Inspect source: FinLBench: A Benchmark for Evaluating Large Language Models on Long-Text Financial Documents →This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.