Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
The FINCH paper unveils a new benchmark for testing AI in real-world finance and accounting tasks, using a dataset built from authentic enterprise documents like Enron spreadsheets and emails. Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 complete less than 40 percent of workflows, struggling most with complex, multi-step, and multimodal tasks. Surprisingly, AI also falters on basic tasks like translation and data entry due to hidden spreadsheet logic.
What it examines
This paper introduces FINCH, a benchmark to test AI agents on real-world finance and accounting workflows using messy, multimodal enterprise data like spreadsheets, emails, and PDFs. It uses authentic business tasks and expert annotation to evaluate how well AI can handle complex, long, and collaborative professional work.
What it concludes
Results show that current AI agents solve less than 40% of real finance workflows, highlighting a big gap in capability. FINCH can help improve AI tools for tasks like financial modeling, reporting, and automation in business, but more research is needed to handle complex, messy, and multi-step enterprise tasks.
Evidence objects
The FINCH paper unveils a new benchmark for AI in finance and accounting, using real-world artifacts like Enron spreadsheets and emails to capture the true complexity and chaos of enterprise workflows.
key_findings bullet 1 · key_findings · validation V0
Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 pass fewer than 40% of FINCH workflows, with performance plummeting on complex, multi-step, and multimodal tasks.
key_findings bullet 2 · key_findings · validation V0
Surprisingly, AI struggles with seemingly simple tasks like translation and data entry due to hidden spreadsheet logic, highlighting a major gap between current AI capabilities and the demands of real enterprise environments.
key_findings bullet 3 · key_findings · validation V0
FINCH presents a novel, large-scale benchmark for AI agents in finance, leveraging authentic, multi-modal workflows from major institutions. Its unique LLM-assisted pipeline and expert annotation address real-world complexity, revealing significant gaps in current AI capabilities. The datasets scale, realism, and practical focus make it compelling and highly original.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
- … employees) and other financial institutions, preserving in-… LLM-assisted discovery with expert annotation: (1) LLM-… for financial decision-making tasks with llm-based agent…
Source row: 834 · abstract type: snippet