Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
Evaluations show agent-security judges can miss attack groups they confidently call safe, limiting automatic decisions.
The paper evaluates System One models Jev, Laya, Decider, and Bespoke Nimble as judges for agent-security decisions, including prompt-injection detection and harmful-request screening. It compares decision accuracy, probability calibration, and selective automation against specialized classifiers and language-model judges. Strong averages can conceal attack groups that models confidently label safe, and policies that meet error limits in validation can exceed them on unseen inputs. A second judge catches some misses but may reject more benign inputs and repeat high-confidence errors.
- High average accuracy can hide attack groups judged safe with high confidence.
- Strict limits on missed attacks allow few inputs to pass automatically.
- Separate allow and block thresholds increase automation mainly by blocking more.
- A second judge can catch misses yet repeat the first model's confident errors.
Full article182 words · extracted from arxiv.org · click to collapse
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models expose typed decisions with probabilities that software can use to allow, block, or escalate inputs, but whether these probabilities support reliable automated security decisions remains unclear. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges, examining decision accuracy, probability calibration, and selective automation. We identify three main findings. (1) Strong overall performance and favorable average calibration can hide systematic failures on particular attack groups, including attacks that models confidently classify as safe. (2) Under strict limits on missed attacks, the evaluated policies allow few inputs automatically, and choosing separate allow and block thresholds increases automation mainly by blocking more inputs. Even policies that meet error limits during validation can exceed them on unseen inputs. (3) A second judge can detect some missed attacks, but it may also reject more benign inputs and repeat the first model's high-confidence errors. These findings show that model accuracy, probability calibration, and the behavior of the resulting decision policy must be evaluated together.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33401