← Back
Evidence source 5185Spot Checked

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

arXiv2025-12-16Paper
Executive summary

The FINCH paper unveils a new benchmark for testing AI in real-world finance and accounting tasks, using a dataset built from authentic enterprise documents like Enron spreadsheets and emails. Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 complete less than 40 percent of workflows, struggling most with complex, multi-step, and multimodal tasks. Surprisingly, AI also falters on basic tasks like translation and data entry due to hidden spreadsheet logic.

What it examines

This paper introduces FINCH, a benchmark to test AI agents on real-world finance and accounting workflows using messy, multimodal enterprise data like spreadsheets, emails, and PDFs. It uses authentic business tasks and expert annotation to evaluate how well AI can handle complex, long, and collaborative professional work.

What it concludes

Results show that current AI agents solve less than 40% of real finance workflows, highlighting a big gap in capability. FINCH can help improve AI tools for tasks like financial modeling, reporting, and automation in business, but more research is needed to handle complex, messy, and multi-step enterprise tasks.

Extracted from this source

Evidence objects

Evidence 421782% extraction confidence
The FINCH paper unveils a new benchmark for AI in finance and accounting, using real-world artifacts like Enron spreadsheets and emails to capture the true complexity and chaos of enterprise workflows.

key_findings bullet 1 · key_findings · validation V0

Evidence 421882% extraction confidence
Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 pass fewer than 40% of FINCH workflows, with performance plummeting on complex, multi-step, and multimodal tasks.

key_findings bullet 2 · key_findings · validation V0

Evidence 421982% extraction confidence
Surprisingly, AI struggles with seemingly simple tasks like translation and data entry due to hidden spreadsheet logic, highlighting a major gap between current AI capabilities and the demands of real enterprise environments.

key_findings bullet 3 · key_findings · validation V0

Evidence 422082% extraction confidence
FINCH presents a novel, large-scale benchmark for AI agents in finance, leveraging authentic, multi-modal workflows from major institutions. Its unique LLM-assisted pipeline and expert annotation address real-world complexity, revealing significant gaps in current AI capabilities. The datasets scale, realism, and practical focus make it compelling and highly original.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

- … employees) and other financial institutions, preserving in-… LLM-assisted discovery with expert annotation: (1) LLM-… for financial decision-making tasks with llm-based agent…

Source row: 834 · abstract type: snippet