SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation
SSP-Bench dynamically generates safety, security, and privacy evals for 24 LLMs, exposing static-benchmark failures like near-zero safety ranking correlation.
The paper introduces SSP-Bench, a dynamic benchmarking framework that generates LLM safety, security, and privacy evaluation instances on demand, ensuring label validity through externally grounded sources, enforcing scope via service-specific validation, and calibrating difficulty with a multi-model steering panel. Benchmark construction is formulated as multi-objective optimization over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench found near-zero correlation in safety rankings due to construct mixing, strong safety-over-refusal coupling, and hidden within-family regressions, showing static benchmarks can misrepresent model behavior.
- On-demand instance generation counters saturation, contamination, and aggregation artifacts of static benchmarks
- Label validity grounded in external sources; difficulty calibrated via multi-model steering panel
- Evaluated 24 models across four safety/security/privacy services
- Reveals near-zero safety ranking correlation and hidden within-family model regressions
Full article151 words · extracted from arxiv.org · click to collapse
Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets often fail under semantically equivalent rephrasings. We introduce SSP-Bench, a dynamic benchmarking framework that generates evaluation instances on demand while preserving domain consistency. The framework ensures label validity through externally grounded sources, enforces scope via service-specific validation, and calibrates difficulty using a multi-model steering panel. Benchmark construction is formulated as a multi-objective optimization problem over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench reveals systematic failures of static evaluation, including near-zero correlation in safety rankings due to construct mixing, strong safety--over-refusal coupling, and hidden within-family regressions. These results show that static benchmarks can misrepresent model behavior, motivating dynamic, deployment-relevant evaluation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.25352