← Back
Finding 4217Emerging EvidenceValidation V0

Despite advances, top AI models such as GPT 5.1 Pro and Claude Sonnet 4.5 pass fewer than 40% of FINCH workflows, with performance plummeting on complex, multi-step, and multimodal tasks.

82%Confidence
1Evidence objects
v1Version
DraftStatus

Evidence trail

Knowledge status

This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.