SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes
FARSIGHT framework finds 80% of 15 academic financial LLM trading schemes fail robustness metrics and 100% exhibit security vulnerabilities.
The paper introduces FARSIGHT, a scheme-level evaluation framework for financial LLM trading agents covering robustness under market turbulence, including flash-crash-like scenarios, and security against attacks on information sources, on agents, and agent-as-attacker behaviors. Applied to 15 representative academic schemes, 80% fail at least one core robustness metric and all 15 show security vulnerabilities. The authors warn the two failure modes are inseparable since small misjudgments can cascade into market-wide crashes that adversaries can deliberately trigger at minimal cost.
- Evaluated 15 academic financial LLM trading schemes
- 80% fail at least one core robustness metric
- 100% exhibit security vulnerabilities across three attack types
- Compromised agents hold direct execution authority over real capital
Full article174 words · extracted from arxiv.org · click to collapse
Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.19705