Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection
A paper analyzes XAI techniques to evaluate Prompt Guard 2's prompt injection detection, finding that explanation methods may reduce the cost of adversarial bypasses.
A paper explores how XAI techniques analyze Prompt Guard 2's ability to distinguish malicious from benign prompts, finding that explanation methods may lower the cost of adversarial bypasses.
70