Finding 4217Emerging EvidenceValidation V0
Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 pass fewer than 40% of FINCH workflows, with performance plummeting on complex, multi-step, and multimodal tasks.
82%Confidence
1Evidence objects
v1Version
DraftStatus
Evidence trail
Supporting82% linkage confidence
Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 pass fewer than 40% of FINCH workflows, with performance plummeting on complex, multi-step, and multimodal tasks.
key_findings bullet 2 · key_findings
Inspect source: Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows →This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.