Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Pretext builds malicious agent skills that evade SkillSpector-style scanners at up to 97 percent.
Pretext is a white-box attacker that iteratively writes agent skills which evade pre-install scanners while still carrying a payload and completing the benign task. Skills are used by agents such as OpenClaw and Claude Code; defenses like NVIDIA's SkillSpector combine static checks with an LLM semantic judge. Moving the payload into natural language defeats static analysis, and framing it as the skill's purpose while splitting instructions across files keeps the LLM score under the block threshold. Across three open-source models, success reached up to 97% against a frozen detector and 77% against a co-adaptive one.
- Pretext moves malicious skill payloads from code into natural-language instructions.
- Legitimate-purpose framing and cross-file splits keep the LLM judge below its block threshold.
- Evasion reaches 97% against a frozen detector and 77% against a co-adaptive one.
- Evaluated defenses include NVIDIA SkillSpector-style scanners for agents such as Claude Code.
Full article156 words · extracted from arxiv.org · click to collapse
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.39607