ZeroHour

Search: “rule-41”

30 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

Who gets to define the rules for AI?

Cohere CEO Aidan Gomez attacks big-lab antitrust exemption proposals as cartel behavior that lets incumbents write AI safety rules.

Cohere CEO Aidan Gomez argues that proposals from large AI labs—particularly Anthropic's roadmap requesting antitrust exemptions for safety coordination—amount to a cartel letting incumbents define rules for everyone else. He draws parallels to the 1975 SEC NRSRO credit-rating designations and the EU's 1985 Motor Vehicle Block Exemption, where safety justifications produced incumbent-protecting market structures. Gomez supports independent review of highly capable AI systems but disputes who writes the standards, who conducts review, and who participates. He also warns AI cyber offense is getting cheaper faster than defenses are improving.

Cyberattack causes a flight delay? Airlines won’t owe you a hotel or meal

A new DOT rule exempts airlines from providing meal vouchers or hotels for cyberattack-caused delays if carriers comply with applicable cybersecurity regulations.

A Department of Transportation rule published in September 2026 adds "cybersecurity attacks" to a list of 10 "not controllable" flight disruption causes, creating a new delay tracking category and relieving compliant airlines of customer service obligations like meal vouchers and hotels. The rule stems from the FAA Reauthorization Act of 2024 and applies only when carriers demonstrate compliance with applicable cybersecurity regulations. Consumer groups reacted cautiously: FlyersRights criticized the lack of public comment, while the National Consumers League saw both certainty benefits and risks from ambiguous wording. The article cites prior aviation incidents including Scattered Spider's airline attacks and the 2024 Collins Aerospace hack that disrupted European flights.

CyberScoop · 4d agoPolicy & legal

Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability

Unit 42 details the Bad Likert Judge multi-turn jailbreak that abuses LLMs' evaluation capability, raising attack success rates over 60% across six frontier models.

Palo Alto Networks Unit 42 describes the Bad Likert Judge technique, a multi-turn jailbreak that asks a target LLM to act as a Likert-scale judge scoring the harmfulness of example responses. The highest-rated example in each scale can carry harmful content, bypassing the model's internal guardrails. Testing across six state-of-the-art text-generation LLMs showed an average attack success rate increase of more than 60% versus plain attack prompts, with tested models anonymized. The technique targets edge cases rather than typical use, and the article positions the work as guidance for defenders on potential jailbreak risks.

Palo Alto Unit 42 · Aug 17, 2026AI safety & security

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

A self-distillation safety framework tunes narrow-boundary refusals in Qwen3-8B, raising target-domain refusal to 84.75% while cutting over-refusal from 15.20% to 5.20%.

The paper formulates narrow-boundary safety, where deployments need refusals within specific topics rather than whole subjects, and proposes an offline self-generated framework with controlled topic generation, escalating retries, and harmful-benign boundary pairs. On political persuasion with Qwen3-8B, the method raised target-domain refusal from 9.47% to 84.75% and cut the mean unsafe-response rate across three broader benchmarks from 26.26% to 0.14%. Verified target-model responses reduced over-refusal from 15.20% to 5.20%, and boundary-pair data cut comply-side over-refusal on held-out pairs from 32.94% to 4.16%. Results show data composition controls the safety-usability trade-off and alignment should be evaluated on both sides of the refusal boundary.

Hugging Face daily papers · 14d agoAI safety & security1

Hiding Prompt Injection in Legal Filing

A judge banned a plaintiff from electronic court filings after hidden prompt-injection text was discovered planted in legal documents.

Bruce Schneier's blog discusses an incident in which hidden prompt-injection instructions were planted inside a legal filing, apparently targeting AI systems that might process court documents. Judge Walter Spader Jr. responded by banning the plaintiff from electronic filings, requiring all future submissions as printed hard copies. Commenters debate whether the tactic could affect future AI-based processing of court records and whether plain-text formats will regain favor.

Schneier on Security · 16d agoAI safety & security in the wild

The EU CRA's Real Question: What Shipped, and When Did You Know?

ActiveState argues the EU CRA's 24-hour ENISA exploit-notification duty, effective September 11, 2026, makes current SBOMs and provenance visibility a legal necessity.

An ActiveState essay warns that the EU Cyber Resilience Act's reporting obligations take effect on September 11, 2026, requiring manufacturers of products with digital elements sold into the EU to notify ENISA within 24 hours of learning a vulnerability is actively exploited, with a fuller report within 72 hours. The law's engineering requirements only apply from December 11, 2027, leaving a visibility-first runway, and Article 13 requires the SBOM to stay current unlike one-time artifacts generated under US Executive Order 14028. The author contrasts the 24-hour notification clock with an industry-average 55 days to remediate high or critical vulnerabilities and recommends automated SBOM regeneration or consuming pre-vetted, attested open source components.

BleepingComputer · 7d agoPolicy & legal

Access Control as Verified Parse Constraints

Researchers verify a class of EverParse validators that correctly enforce access-control policies, deploying a machine-checked enforcement gate on seL4.

The paper targets enforcement-code bugs in commercial security gateways by proving that forward-only, backtrack-free EverParse validators are verified recognizers for a bounded finite-state class that includes access-control decision functions with fixed-offset fields and bounded disjunction. Encoding a bounded policy language into a fixed-size byte buffer allows an SMT solver to verify the enforcement code once, covering all byte values, policies, requests, and sessions. Editing rule content over a fixed endpoint set requires no new proof, while adding endpoints reruns the toolchain. A deployment on the seL4 microkernel ensures every request passes through the gate and unverified components cannot corrupt the enforcement chain.

arXiv cs.CR · 5d agoResearch

Cyber threats nudge Trump to sign executive order on foreign equipment in U.S. energy infrastructure

Trump signed an executive order declaring an emergency to bar foreign bulk-power equipment deemed a national security cyber risk.

The executive order, 'Declaring a National Energy Emergency to Secure the United States Bulk-Power System,' prohibits acquiring, importing, transferring, or installing foreign-produced bulk-power equipment and software deemed risky, citing fears of digital backdoors in Chinese-made grid gear. China supplies roughly 85% of solar supply chain capacity and is a major transformer manufacturer. The Energy Department has 120 days to develop implementing rules; the order revives a 2020 Trump-era measure the Biden administration had suspended after utilities found compliance difficult.

CyberScoop · 20d agoPolicy & legal1

Closing the Gap Between Detection and Protection with AI-Assisted Custom Rules

Akamai describes using AI-assisted custom rules to close the gap between threat detection and active protection in security operations.

Akamai published a blog post on AI-assisted custom rules intended to close the gap between detecting threats and enforcing protections. The post appears to be a vendor capability discussion for security operations teams. No article text was available beyond the title, so further technical details are limited.

Akamai Blog · 23d agoTools1

Managing the cyber risk of agentic AI

UK NCSC guidance recommends safeguards, sandboxing, and active oversight to manage cyber risks of autonomous agentic AI systems.

The UK National Cyber Security Centre published guidance on managing the cyber risk of agentic AI systems. It recommends safeguards, sandboxing, and active human oversight to limit unintended autonomous activity while realizing the benefits of these systems. The publication is official national guidance for organizations deploying agentic AI.

NCSC UK · 27d agoAdvisory

Severity Is Not a Strategy: What CISA BOD 26-04 Means for the Future of Federal Software Security

CISA's BOD 26-04 replaces severity-based federal patching with risk-based remediation deadlines of 3, 14, or 60 days.

CISA's Binding Operational Directive 26-04, released June 10, 2026, replaces BOD 19-02 and BOD 22-01 for Federal Civilian Executive Branch agencies and shifts remediation prioritization from CVSS scores to risk context. Agencies assess four factors: public exposure, KEV listing, exploit automatability, and whether exploitation grants partial or total asset control, resulting in 3-, 14-, or 60-day remediation windows or next-upgrade fixes. In CISA's first review at a large civilian agency, only 1% of vulnerabilities required three-day remediation while over 60% could wait for future system upgrades. The directive also requires forensic analysis when exploitation is suspected, and Checkmarx argues the same risk-based logic must extend upstream into software development and SBOM-driven exposure management.

Checkmarx · 7d agoPolicy & legal

ChatGPT and Reddit now face EU's toughest online safety rules

ChatGPT and Reddit now fall under the EU's toughest online safety rules, adding new regulatory burdens after rapid growth.

Ars Technica reports that ChatGPT and Reddit are now subject to the European Union's strictest online safety rules, following their explosive user growth. This brings the AI chatbot and the social platform under heightened EU oversight and compliance obligations. The move signals that fast-scaling AI consumer products face the same regulatory scrutiny as major online platforms in the EU.

Ars Technica · AI · 16d agoAI policy

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Multiverse Computing's Hugging Face post argues language models should refuse only the relevant subset of a topic instead of over-refusing whole subjects.

A Hugging Face blog post by Multiverse Computing examines refusal granularity in language models, arguing models should refuse the relevant subset of a topic rather than the entire topic. No full article text was available for additional technical detail.

Hugging Face Blog · 8d agoAI safety & security

FTC Withdraws Obsolete Policy Statement

The FTC rescinded its 2021 policy statement that applied the Health Breach Notification Rule to health apps and connected devices collecting consumer health data.

The Federal Trade Commission formally rescinded its 2021 Policy Statement on Breaches by Health Apps and Other Connected Devices. The statement had purported to apply the FTC's Health Breach Notification Rule to health apps and connected devices that collect consumer health information. The Commission considers the statement obsolete following its 2024 update to the Health Breach Notification Rule.

DataBreaches.net · 6d agoPolicy & legal

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

LexFlip releases 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving tokens, exposing weaknesses in embedding-based meaning preservation metrics.

LexFlip provides 373 minimal perturbations of Quebec statutory French that reverse legal force while preserving 0.93 of tokens, creating dissociation items that break monotone token-overlap metric validation. The seven embedding and BERTScore metrics tested register only 0.022-0.039 of their identical-to-unrelated range on these edits, versus 0.670 for bidirectional NLI. Against FrJudge, with a measured human ceiling of r=0.597, a bare length feature outscores every semantic metric tested.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

Microsoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety Constraints

Microsoft AI's draft Humanist AI Code of Conduct blocks MAI models from producing exploit code and constrains autonomous agent behavior.

The draft code sets 'Absolute Constraints' preventing MAI models from generating working exploit code, attack tooling, or intrusion guidance, while permitting authorized defensive work such as vulnerability discovery and malware analysis. A 'Chain of Command' rule means tool outputs, file contents, and webpages carry no authority over model behavior, countering injected instructions. Microsoft opened a six-week public consultation; a revised version will guide 2027 model development, and current MAI Models were not trained on the document.

SecurityWeek · 1d agoAI safety & security1

Who's governing your AI? A trust framework for enterprise agents and models

DigiCert pitches AI Trust framework using PKI, DNS policy records and workload identity to govern shadow AI agents across enterprises.

The Register-sponsored piece outlines DigiCert's AI Trust framework for governing AI agents, built on PKI, DNS, and attestation, citing IBM's 2026 Cost of a Data Breach report that 68% of organizations lack AI governance or shadow AI detection. The approach treats agent identity as workload identity aligned with IETF WIMSE, NIST CSF 2.0, and SPIFFE/SPIRE, using short-lived credentials instead of static API keys. DigiCert also proposes DMARC-style DNS agent policy records and an AI Agent Passport cryptographically binding agent identity to approved operations, with a unified kill switch.

The Register · Security · 1d agoAI safety & security1

New AI Attack Hides Malicious Instructions in Normal-Looking Text to Evade Safety Filters

Check Point researchers show crafted prose hides policy-violating instructions that bypass all tested LLM gatekeepers, including GPT-4o mini and Llama Guard 3.

A new prompt-crafting technique embeds malicious payloads inside grammatical, natural-looking text without Base64, invisible Unicode, or obvious encodings, defeating lightweight pre-screening gatekeepers. In testing, all four evaluated gatekeeper models—gpt-4o-mini-2024-07-18, gpt-oss-safeguard:20b, claude-3-haiku-20240307, and llama-guard3:8b—classified the crafted wrappers as safe at a 100% bypass rate across 23 obfuscated prompts. GPT-5 Thinking in high-reasoning mode recovered and acted on the hidden instruction in 17 of 18 tests (~94.4%), often spending over a minute and multiple Python executions. Researchers recommend paraphrasing untrusted input, hardening gatekeeper policies, and applying defense-in-depth controls for agentic deployments.

GBHackers · 5d agoAI safety & security 2 sources

Risky Bulletin: White House lets private companies carry out offensive cyber ops

A White House memo directs DHS to create a program letting vetted private companies conduct US-government-directed offensive cyber operations against cybercrime.

A presidential memo tasks the DHS National Coordination Center with building a program, under DOJ and DHS oversight, through which private-sector companies can conduct offensive cyber operations against large-scale cybercrime organizations. Requirements include secure facilities, vetted personnel, a $1 million escrow for damages, and written approvals co-signed by DHS and DOJ executive directors. The program must launch within 60 days, around October 11, expanding a March executive order targeting scam compounds, ransomware, and other large-scale cybercrime.

Risky Business News · Aug 14, 2026Policy & legal

USN-8736-2: Perl vulnerabilities

Ubuntu issued USN-8736-2 fixing two Perl regex flaws that could cause out-of-bounds heap access, denial of service, or security bypass on 24.04 LTS.

USN-8736-2 backports the Perl fixes from USN-8736-1 to Ubuntu 24.04 LTS. CVE-2026-15534 involves out-of-bounds heap reads or writes when regular expressions handle large inputs, potentially causing denial of service or arbitrary code execution. CVE-2026-19487 causes incorrect regex matching with alternative branches, allowing security restrictions to be bypassed.

Getting ahead of ‘harvest-now-decrypt-later’: Post-quantum cryptography planning

Opinion piece urges organizations to begin post-quantum cryptography migration now, citing harvest-now-decrypt-later risk and NIST deadlines.

CSO Online outlines why harvest-now-decrypt-later makes long-lived sensitive data a current risk even before quantum computers exist. It cites NIST IR 8547 timelines deprecating RSA-2048 and ECC P-256 by 2030 and removing them by 2035, finalized FIPS standards ML-KEM, ML-DSA, and SLH-DSA, upcoming FN-DSA (FIPS 206), NSA requirements for national security systems from 2027, and UK NCSC phased guidance through 2035. The author recommends cryptographic discovery, crypto-agility, and prioritizing long-confidentiality data and TLS endpoints.

CSO Online · 6d agoResearch

17 draft Cyber Resilience Act standards are open for comment

ETSI publishes 17 draft harmonised standards detailing EU Cyber Resilience Act compliance, open for comment until between mid-September and mid-November 2026.

Seventeen draft standards covering the higher-risk tier of products with digital elements, including password managers, antivirus software, connected toys and wearables, are open for comment. Following a Harmonised Standard grants manufacturers the presumption of conformity with the Cyber Resilience Act, whose obligations apply through the end of 2027 to importers, distributors, service providers and developers. The drafts went to 41 member organisations plus societal partners ANEC, ECOS, ETUC and SBS, with closing dates varying by vertical.

Help Net Security · Aug 14, 2026Policy & legal

Invisible AI Prompts Trigger Court Sanctions

A Connecticut litigant hid white-font prompt injections in court filings to sway AI systems; the judge sanctioned him by revoking e-filing privileges.

A self-represented plaintiff hid prompt injection instructions in 3-point white text within court filings, telling any AI model reading the documents to agree with his filings and grant him relief. The judge called it serious litigation abuse and sanctioned him by revoking electronic filing privileges. It is reportedly the first documented prompt injection attack against a US court and the first sanction for attempting one.

Security Affairs · Aug 17, 2026AI safety & security in the wild

Optimizing Credential Blast Radius Through Trust Boundaries and Delegation Under Post-Quantum Authentication Costs

Academic paper models credential blast radius optimization across trust domains under post-quantum latency costs, cutting expected impact by up to 36%.

The paper formulates the joint selection of trust domains and credential-derivation structures under policy and latency constraints as an NP-hard optimization problem, showing the scalarized two-domain direct-issuance case reduces to a weighted minimum cut. In 195 of 230 exhaustive synthetic comparisons, joint optimization produced lower credential blast radius than choosing boundaries first, especially under chained delegation. A trace-derived replay using measured post-quantum authentication costs found the best design reduced expected impact by up to 36% relative to a single domain within the latency budget.

arXiv cs.CR · 12d agoResearch

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

A review paper frames on-policy self-distillation collapse as governed by three levers: token weighting, privileged information, and guidance decay.

The paper critically reviews On-Policy Self-Distillation (OPSD), where a language model trains on its own generations scored token-by-token by a teacher conditioned on privileged information such as reference solutions or environment feedback. It identifies collapse, the progressive narrowing of producible reasoning paths, as the dominant failure mode and analyzes it through three levers: signal weighting, the nature of privileged information, and teacher dynamics. The review is restricted to mathematical reasoning, reports no new experiments, and offers a shared vocabulary separating settled findings from disputed ones.

Hugging Face daily papers · 22d agoAI research

From Intent to Execution Grant: An Execution-Boundary Conformance Profile for High-Risk AI Actions

Researchers specify EBL-Core, an execution-boundary conformance profile binding AI agent intents, policies, and evidence into verifiable execution grants, validated with bounded tests.

The paper defines EBL-Core, a conformance profile deciding whether one fully materialized AI-generated candidate action may receive action-scoped execution authority. It binds a structured intent object, Root and Operational Policies, typed evidence, and a verifiable Decision Derivation through an Execution Release Contract, with lifecycle rules for Redemption and Revocation. Evaluation included 34 static vectors, 15 lifecycle checks, and 100 trials of 32 concurrent Redemption attempts yielding exactly one winner per trial. The authors state these bounded results demonstrate executability of the specified subset, not production readiness or complete mediation.

arXiv cs.CR · 6d agoAI safety & security1

Governing Bring Your Own AI: A Parameterized Maturity Model

Researchers propose a parameterized governance model and maturity ladder for Bring Your Own AI, finding data exposure and compliance dominate BYOAI risks.

The paper studies Bring Your Own AI (BYOAI), where employees use personal generative AI accounts such as ChatGPT, Gemini, and Claude outside enterprise identity and security controls. Drawing on a curated corpus of 30 records (24 studies and 6 framework documents), the authors build a risk taxonomy, a five-level governance maturity ladder, and a parameterized model linking control-layer coverage to residual risk. Findings highlight data exposure and compliance as the most prominent risks, inconsistent framework engagement, and evidence that layered technical controls reduce modeled exfiltration risk more than prohibition-based approaches.

arXiv cs.CR · 12d agoResearch

CVE-2026-86304: MojoX::Authentication versions before 0.006 for Perl allow SAML authentication bypass because parse_assertion builds Net::SAML2::Binding::POST without a trust anchor

MojoX::Authentication before 0.006 for Perl allows SAML authentication bypass because parse_assertion builds Net::SAML2::Binding::POST without a trust anchor (CVE-2026-86304).

CVE-2026-86304 affects MojoX::Authentication versions before 0.006 for Perl. The parse_assertion function builds Net::SAML2::Binding::POST without a trust anchor, so SAML assertions are not validated against a trusted signing key, enabling authentication bypass. The flaw is fixed in version 0.006 of the module.

oss-security · 9d agoVulnerabilityCVE-2026-86304

EU Cyber Resilience Act to Enforce New Reporting Requirements

EU Cyber Resilience Act reporting obligations begin Friday, requiring businesses to notify serious product security incidents within 24 hours.

The EU Cyber Resilience Act's new reporting requirements take effect starting Friday. Businesses operating in the EU will have 24 hours to notify the government whenever they discover serious product security incidents.

Dark Reading · 6d agoPolicy & legal

TCRF taken offline by DDoS attack after Claude user ban

The Cutting Room Floor game wiki was taken offline by a DDoS attack after a user leveraging Claude was banned.

The Cutting Room Floor (TCRF), a wiki documenting unused video game content, was knocked offline by a distributed denial-of-service attack. The attack reportedly followed moderation action banning a user who was using Anthropic's Claude. The incident highlights friction between community sites and AI-assisted users and tools.

Lobsters · security · 18d agoAI safety & security in the wild