FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
Researchers have launched FINVAULT, the first security benchmark for financial AI agents, exposing serious vulnerabilities in large language model-powered systems. Tests show even top models can be breached in over 20 percent of cases, with some failing up to 85 percent, especially in insurance. FINVAULT simulates 31 real regulatory cases and 107 vulnerabilities, revealing that semantic attacks like role-playing are more effective than technical ones. Current defenses miss attacks or trigger too many false alarms.
What it examines
This paper introduces FINVAULT, a benchmark to test the safety of financial AI agents using real-world scenarios, databases, and compliance rules. It aims to reveal security risks in financial environments that current evaluations miss, focusing on agents' actions and vulnerabilities in realistic, execution-grounded workflows.
What it concludes
FINVAULT shows that financial AI agents are still vulnerable to attacks, especially those using clever language tricks. Existing defenses are not strong enough. This research can help improve security for financial AI systems used in banking, insurance, and trading, and guides future work on safer, more reliable financial agents.
Evidence objects
FINVAULT, the first execution-grounded security benchmark for financial AI agents, reveals that leading LLM-powered financial models are alarmingly vulnerable, with compromise rates ranging from 20% to a staggering 85% in complex domains.
key_findings bullet 1 · key_findings · validation V0
By simulating 31 real regulatory cases and 107 vulnerabilities, FINVAULT enables concrete measurement of security failures, showing semantic attacks like role-playing and emotional manipulation are far more effective than technical exploits.
key_findings bullet 2 · key_findings · validation V0
The study finds current defenses are impractical, missing many attacks or causing excessive false alarms, and notes FINVAULTs sandboxed tests, while realistic, cannot fully capture the complexity of live financial systems.
key_findings bullet 3 · key_findings · validation V0
FINVAULT presents the first execution-grounded security benchmark for LLM-powered financial agents, uniquely addressing real-world compliance and safety gaps. Featuring 31 regulatory scenarios, 107 vulnerabilities, and 963 test cases, its novel methodology enables rigorous, realistic assessment, offering significant impact for both academic research and industry deployment in financial AI agent safety.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
Abstract: Financial agents powered by large language models (LLMs) are increasingly deployed for investment analysis, risk assessment, and automated decision-making, where their abilities to plan, invoke tools, and manipulate mutable state introduce new security risks in high-stakes and highly regulated financial environments. However, existing safety evaluations largely focus on language-model-level conten… ▽ More Financial agents powered by large language models (LLMs) are increasingly deployed for investment analysis, risk assessment, and automated decision-making, where their abilities to plan, invoke tools, and manipulate mutable state introduce new security risks in high-stakes and highly regulated financial environments. However, existing safety evaluations largely focus on language-model-level content compliance or abstract agent settings, failing to capture execution-grounded risks arising from real operational workflows and state-changing actions. To bridge this gap, we propose FinVault, the first execution-grounded security benchmark for financial agents, comprising 31 regulatory case-driven sandbox scenarios with state-writable databases and explicit compliance constraints, together with 107 real-world vulnerabilities and 963 test cases that systematically cover prompt injection, jailbreaking, financially adapted attacks, as well as benign inputs for false-positive evaluation. Experimental results reveal that existing defense mechanisms remain ineffective in realistic financial agent settings, with average attack success rates (ASR) still reaching up to 50.0\% on state-of-the-art models and remaining non-negligible even for the most robust systems (ASR 6.7\%), highlighting the limited transferability of current safety designs and the need for stronger financial-specific defenses. Our code can be found at https://github.com/aifinlab/FinVault. △ Less
Source row: 878 · abstract type: unknown