Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
Curriculum RL red-teams frontier LLMs for prompt injection, hitting 93.8% ASR against GPT-5.6-Luna on AgentDyn.
Researchers propose curriculum reinforcement learning to red-team frontier LLMs for prompt injection, avoiding a cold start where every attempt scores zero reward. The attacker is trained on a sequence of increasingly robust targets and must partially succeed at each stage. On AgentDyn, the method reached ASR@10 of 93.8% against GPT-5.6-Luna and 45.0% against GPT-5.6-Terra, while RL-Hammer and PISmith scored 0%. An attacker trained on GPT-5.6-Terra also transferred to six other frontier models, including GPT-6-Luna; code is released as PIForge.
61