ZeroHour

Search: “reliability”

5,506 stories

When LLM judges agree, should we believe them?

Amazon ICML paper uses Ising models to correct correlated LLM-judge votes, beating accuracy-weighted panels by 9-14%.

Amazon Science describes an ICML paper, "Dependence-aware label aggregation for LLM-as-a-judge via Ising models," addressing how correlated judge outputs inflate majority-vote confidence. The unsupervised method models pairwise dependence between judges, learning both reliability and similarity without human reference labels. Tested on relevance, toxicity, and summarization tasks with 10 judge models at temperature zero, it outperformed accuracy-weighted voting by 9% to 14%.

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Audit of eight cybersecurity LLM benchmarks shows evaluation pipeline choices can swing scores by over 80 points and reshuffle most model rankings.

Researchers modeled eight cybersecurity benchmarks as configurable measurement pipelines and audited 10 proprietary, open-weight, and cybersecurity-specialized LLMs. They identified 15 systematic failure modes and showed a single pipeline choice can change a model's score by more than 80 percentage points and alter rankings; semantically similar task pairs rank the same models differently. Under a standardized harness, nine of 10 models shifted at least three ranks on at least one benchmark, motivating pipeline-aware auditing for reliable model evaluation.

arXiv cs.CR · 8d agoAI research1

Why Is SHAP Not a Reliable Standalone Explanation Framework for Malware Detection?

arXiv paper shows SHAP gives unreliable standalone explanations for malware detection, with attribution dilution and sign reversal in dependent PE feature spaces.

The paper argues SHAP's formal guarantees are insufficient for reliable malware interpretation because the explained feature-coalition game is fixed only by analyst choices, not by malware behavior in the data. In static Portable Executable feature spaces, dependent feature groups cause conditional SHAP to dilute credit by a factor of 1/m across redundant features, attribute importance to features the model never uses, and even reverse attribution signs; interventional SHAP queries off-manifold coalitions no real executable exhibits. Experiments on EMBER-2018, EMBER-2024, and BODMAS with fixed LightGBM and XGBoost detectors confirm these effects. The authors position SHAP as a limited diagnostic requiring explicit data-distribution statements and domain validation, not a standalone explanation framework.

arXiv cs.CR · 12d agoResearch1

Bipartisan Senate bill aims to prepare energy sector for Q

Bipartisan Senate bill would direct FERC to factor quantum computing threats and post-quantum cryptography into US electric grid cybersecurity reliability standards.

The Quantum Grid Utility Assurance and Resilient Defense (Quantum-GUARD) Act, introduced by Senators Mike Rounds and Chris Coons, would require FERC to consider quantum computing threats when reviewing electric reliability standards and to explore post-quantum cryptography use in both IT and OT systems, plus a technical sandbox to study quantum impacts. It aligns with NIST's post-quantum algorithm work, and a June executive order moved the federal PQC migration deadline from 2035 to 2030. Industry experts noted the hard part is upgrading infrastructure such as SCADA communications and software update integrity ahead of those deadlines.

CyberScoop · 22d agoPolicy & legal