A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
A positive safe-response centroid fails as a safety direction; an explicit safe-minus-unsafe reference scores far better.
Researchers audit a sleeper-agent detector that scores response safety by cosine similarity to the mean embedding of known-safe responses. On human-labeled PKU-SafeRLHF and Aegis, the safe prototype reaches ROC-AUC 0.457-0.545, sometimes below chance, while a safe-minus-unsafe reference reaches 0.588-0.738; on a jury-labeled control the prototype inverts to 0.358-0.405 and the reference reaches 0.754-0.793. At a 5% false-safe threshold the reference accepts more truly safe responses on PKU-SafeRLHF and Aegis, but not reliably on BeaverTails. Prompt-label composition can inflate uncontrolled evaluations, and the finding applies only to a raw positive centroid.
- Safe prototype ROC-AUC is 0.457-0.545 on human-labeled corpora.
- Safe-minus-unsafe reference reaches 0.588-0.793 depending on corpus.
- Eighty to 634 labeled unsafe responses recover most of the ranking.
- Prompt-only ablations show label composition can inflate scores.
- Result is limited to raw positive centroids, not all one-class methods.
Full article226 words · extracted from arxiv.org · click to collapse
Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.01801