← Back
Finding 4247Emerging EvidenceValidation V0

Commercial LLMs consistently outperform open-source models on FinLBench tasks, but all models struggle with hallucinationfabricating information, especially when confronted with trap questions, exposing a critical and surprising weakness in current AI systems.

78%Confidence
1Evidence objects
v1Version
DraftStatus

Evidence trail

Knowledge status

This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.