An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
An open dashboard scores 18 models on EU GPAI systemic risks, showing averages hide 14–37 point worst-case drops.
Researchers introduce the Systemic Risk Index, an open pipeline and dashboard that traces model risk scores to public benchmark evidence under the EU GPAI Code of Practice. It groups 19 benchmarks into CBRN, cyber offense, harmful manipulation, and loss of control, and tests models with harm-preserving perturbations and simulated deployment contexts. Across 18 models, worst-case aggregation lowers scores by 14 to 37 points compared with averages. LLM judges agree with human graders at kappa 0.78–0.82, and 83 percent of audited transformations preserve the original harm.
- Organizes 19 public benchmarks into four EU GPAI systemic-risk categories.
- Worst-case aggregation lowers scores by 14 to 37 points versus averages.
- LLM-judge agreement with humans matches human-human kappa of 0.78 to 0.82.
- Blind audit finds 83 percent of sampled transformations preserve original harm.
Full article195 words · extracted from arxiv.org · click to collapse
Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($κ= 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.28335