Chaining Skills to Hijack LLM Agents
APEX skill chains hijack LLM agents into attacker-chosen actions in 74% of attempts.
The paper introduces APEX, which builds adversarial skill chains so an upstream skill plants a false user-approval record that a downstream skill uses to trigger an attacker-selected action. Across four action families and six models on SkillsBench, chains succeeded in 512 of 690 attempts (74.2%). On GPT-5.4, full chains hit 84.3% versus 17.4% for a single merged skill. A prompting defense cut GPT-5.4 success to 59.1% but also dropped the benign verifier pass rate from 86.7% to 56.3%.
- APEX chains false approval records across agent skills
- 74.2% success across 690 attempts on six models
- GPT-5.4 full-chain success 84.3% versus 17.4% merged
- Prompt defense cuts attacks but hurts benign task performance
Full article216 words · extracted from arxiv.org · click to collapse
LLM agents use skills to improve performance on specialized tasks. To complete a user request, an agent may invoke several skills in sequence, allowing information produced under one skill to guide the next. Because skills may come from open-source repositories, this handoff can also carry attacker-controlled claims into later decisions. In this paper, we introduce APEX, which constructs and refines adversarial skill chains tailored to a user task and an attacker-selected action. The key insight is that an agent-written record of genuine task progress can carry a false claim of user approval across skills: an upstream skill induces the agent to create the record, and a downstream skill uses it to direct the attacker-selected action. Across four targeted-action families and six models on SkillsBench, the chains induce the selected action in 512 of 690 attempts (74.2%). On GPT-5.4, the full chain succeeds in 84.3% of attempts, compared with 17.4% when the workflow is merged into one skill. We further evaluate a prompting defense that asks the agent to check skill-produced files against the original request. On GPT-5.4, it lowers targeted-action success from 84.3% to 59.1%, while the verifier test-pass rate across 72 benign native-skill tasks falls from 86.7% to 56.3%. These results highlight the need for defenses that prevent attacker-directed actions while preserving legitimate task performance.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.01564