ZeroHour

Search: “refusal”

2,013 stories

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

Study shows training LLMs on refusal rationales instead of boilerplate refusal statements reduces false refusals while maintaining safety performance.

The paper decomposes safety-tuning responses into a boilerplate refusal statement and an explanatory rationale, finding that refusal statements push models to rely on superficial cues and misjudge benign queries as harmful. Training solely on rationales reduces false refusals while maintaining comparable safety performance, and the benefits carry over to in-context learning configurations and remain compatible with inference-time mitigations. The results argue for precisely curated, fine-grained safety supervision datasets when aligning LLMs.

Hugging Face daily papers · 12d agoAI safety & security1

dealignai/GLM-5.3-CYBERSECURITY-FP8 — new model trending #13 on Hugging Face

dealignai releases GLM-5.3-CYBERSECURITY-FP8, a 753B MoE weight-modified variant cutting refusals on offensive-security prompts, trending #13.

dealignai released GLM-5.3-CYBERSECURITY-FP8 on Hugging Face, a cybersecurity-domain 'crack' of the 753B-parameter GLM-5.3 MoE model, currently trending #13. The release directly edits bf16 residual writers, keeps FP8 routed experts, and serves with stock vLLM on 8x H200 GPUs with 131k context. HarmBench-320 evaluations show 80-84% direct harm compliance and 89% cyber-offense compliance, while MMLU rose 1.07 points to 86.65%. Copyright-verbatim reproduction remains a known soft-refusal limitation, with an UNCENSORED sibling variant offered.

Hugging Face trending models · 16d agoModel release

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

A self-distillation safety framework tunes narrow-boundary refusals in Qwen3-8B, raising target-domain refusal to 84.75% while cutting over-refusal from 15.20% to 5.20%.

The paper formulates narrow-boundary safety, where deployments need refusals within specific topics rather than whole subjects, and proposes an offline self-generated framework with controlled topic generation, escalating retries, and harmful-benign boundary pairs. On political persuasion with Qwen3-8B, the method raised target-domain refusal from 9.47% to 84.75% and cut the mean unsafe-response rate across three broader benchmarks from 26.26% to 0.14%. Verified target-model responses reduced over-refusal from 15.20% to 5.20%, and boundary-pair data cut comply-side over-refusal on held-out pairs from 32.94% to 4.16%. Results show data composition controls the safety-usability trade-off and alignment should be evaluated on both sides of the refusal boundary.

Hugging Face daily papers · 13d agoAI safety & security1

Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions

A linear hidden-state direction encodes question impossibility in 1.7B-70B LLMs, but misalignment with the safety-refusal pathway explains why models answer unanswerable questions.

The study examines why instruction-tuned LLMs from 1.7B to 70B parameters answer structurally unanswerable math and code questions instead of abstaining. A single linear direction in the hidden state separates answerable from impossible prompts, showing models represent impossibility before generation, but this direction is nearly orthogonal to the canonical safety-refusal direction. Generation-time steering along the recognition direction changes invalidity-aware behavior dose-responsively, and the geometry is present even at the pretraining endpoint, indicating a routing failure rather than an encoding failure.

Hugging Face daily papers · 18d agoAI safety & security

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

Researchers show directional ablation breaks refusal in GLM-5.3-Flash, a 320B-parameter MoE, cutting refusal by 41–89 points across seven benchmarks.

The study extends directional ablation, a white-box attack that removes an aligned LLM's refusal behavior, from dense models up to ~70B parameters to GLM-5.3-Flash, a 320B-parameter mixture-of-experts model with 288 routed experts, four-wide hyper-connection residual, and block-FP8 quantization. Editing attention, dense, and routed-expert writers jointly removes 0.776 of refusal, with 74% of the effect existing only under the joint intervention; the conventional module-name-based recipe reaches only 0.066 and fails silently on MoE architectures. The attack yields 41–89 percentage-point reductions in refusal across seven harmful benchmarks with no detected capability change, and a category-concentrated refusal residue survives all edits at ranks 1 to 12.

arXiv cs.CR · 6d agoAI safety & security

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Multiverse Computing's Hugging Face post argues language models should refuse only the relevant subset of a topic instead of over-refusing whole subjects.

A Hugging Face blog post by Multiverse Computing examines refusal granularity in language models, arguing models should refuse the relevant subset of a topic rather than the entire topic. No full article text was available for additional technical detail.

Hugging Face Blog · 7d agoAI safety & security

Rhysida Publishes Berlin Government Data After €2m Extortion Demand Refused

Rhysida published 5.7 TB of Berlin state government data, including sensitive CBRN emergency plans, after the state refused a €2m ransom demand.

The Rhysida ransomware gang leaked roughly 5.7 TB — about 1.4 million files — stolen from Berlin's state network after the government declined to pay 30 bitcoins (about €2m) by the September 4 deadline. The dump reportedly includes sensitive state emergency plans for terrorist attacks and CBRN disaster scenarios in a folder titled 'AG CBRN-Rahmenplanung', plus personnel files, absence lists, payroll data, and home addresses, potentially affecting tens of thousands of people. Berlin says there are no indications the state network remains compromised and will notify affected individuals on a risk-based basis after forensic analysis.

Infosecurity Magazine · 8d agoRansomware

Berlin Refuses to Pay Hackers Who Stole Data From the City's State Network

Berlin refuses ransom demands after Rhysida claimed stealing 5.79 TB from the city-state network, including mobility department data on 12,076 individuals.

Berlin's state government confirmed an extortion attempt following the August compromise of its state administrative network and said it will not pay. Forensics found data exfiltrated from the Senate Department for Mobility, Transport, Climate Protection and Environment between August 7 and 12, 2026; Rhysida's leak site claims 5.79 TB, 1.44 million files, and personal data on 12,076 individuals. State criminal police, prosecutors, and federal authorities are investigating, and officials say the September 20 Abgeordnetenhaus election environment remains unaffected. Separately, Manchester Airports Group confirmed theft of customer data including emails, phone numbers, and vehicle registrations across three airports.

The Hacker News · 18d agoRansomware in the wildCVE-2020-1472

Introducing the CyberAgents Exchange AI Inspector: Rigorous review for community-built AI

Tenable and OpenAI launch the CyberAgents Exchange AI Inspector to security-review community-submitted AI agents, MCP servers, and skills using GPT Cyber models.

Tenable and OpenAI announced the CyberAgents Exchange AI Inspector, unveiled at OpenAI's "Intelligence at Work: Cyber Summit," to vet community-submitted AI agents, skills, MCP servers, and multi-agent playbooks in the CyberAgents Exchange registry. The process combines Tenable One AI Exposure scanning, OpenAI GPT Cyber model assessment, and human review, with reviews anchored to specific Git commits. The registry launched in August and hosts over 100 AI listings; the Inspector is expected to be available in September and has already detected prompt injection implemented via invisible Unicode tag characters in a SKILL.md file.

Tenable Blog · 6d agoTools

Berlin Ransomware Leak Exposes State Secrets

Rhysida ransomware leaked 5.79 TB of Berlin state government data on the dark web after the city refused a 30 Bitcoin ransom.

The Rhysida ransomware group published 5.79 TB (about 1.44 million files) of Berlin state administration data on its leak site on August 28, 2026, after the city-state refused a 30 Bitcoin ransom. The claimed dataset includes personal data on 12,076 individuals, over 5,000 personnel files, payroll records, plaintext credentials for systems like PAYONE and Z_ADMIN, Bundesrat committee protocols, and documents allegedly containing state secrets. Leaked material reportedly covers national defense emergency plans, federal secret communication channels, a CBRN threat planning folder, and vulnerability analyses of Berlin's water supply. Berlin's government has activated a central crisis unit to review the leaked data and notify affected citizens and businesses.

Security Affairs · 8d agoRansomware