ZeroHour

Search: “defender-experts”

30 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

The AI security question leaders should be asking instead

Gremlin security officer Frederic Bull argues AI has eroded the attacker-defender skill asymmetry while least-privilege controls remain essential for securing AI agents.

In a Help Net Security interview, Gremlin Security Officer Frederic Bull says AI has narrowed the expertise gap between attackers and defenders, enabling faster exploit discovery even by less-skilled actors. His team processed roughly nine times more vulnerabilities in the past year with unchanged staffing using LLM-based tooling, cutting time-to-remediate by about 5%. He argues least privilege, session-based RBAC via OIDC/OBO, and human-in-the-loop oversight remain the bedrock defenses for AI agents, and that hiring should favor engineers able to catch confidently wrong AI output.

Help Net Security · 4h agoIndustry

The Defender’s Window

OpenAI details how AI is reshaping offensive and defensive cybersecurity and outlines steps security teams should take now.

OpenAI published 'The Defender's Window,' discussing how AI is changing cybersecurity for both attackers and defenders. The piece describes how OpenAI is strengthening its own defenses and offers guidance on what security teams can do now to adapt, functioning largely as corporate thought leadership rather than an incident or research disclosure.

OpenAI News · Aug 17, 2026AI industry

“Sorry, I can’t help with that”: How your guardrails might become the attacker’s best friend

Cisco Talos's David Bianco argues AI guardrail customization requires operational sovereignty so defenders retain the advantage over attackers.

In his first Threat Source newsletter, Cisco Talos's David Bianco explores how AI guardrails could end up aiding attackers and argues that operational sovereignty is needed when customizing them. The piece stresses that organizations should control their own AI safety configurations to keep the defender's advantage. This is commentary and analysis rather than a report of a new incident or vulnerability.

Cisco Talos · 20d agoAI safety & security

Detect and disrupt AI-themed attacks with Microsoft Defender

Microsoft Threat Intelligence reports criminal campaigns impersonating ChatGPT, Copilot, Claude, and DeepSeek in phishing, AiTM, and malvertising attacks reaching 100,000 emails daily.

Microsoft Threat Intelligence observed a growing set of campaigns that abuse trust in popular AI brands: a ChatGPT-themed phishing campaign sent up to 100,000 emails in one day to steal payment card data, and a Claude-themed campaign used adversary-in-the-middle techniques to harvest credentials and access tokens. Other campaigns included malvertising for a fake AI Windows plugin delivering the Vidar stealer and fraudulent DeepSeek installers distributed via GitHub. Initial access broker Storm-3075 used AI-themed malvertising to distribute payloads for multiple downstream actors, and Microsoft notes the AI services themselves were not compromised. Microsoft also details Defender protections such as Safe Links, Safe Attachments, and attack disruption against these multi-stage lures.

Microsoft Security Blog · 6d agoPhishing & fraud in the wild 2 sources1

OpenAI Builds ‘Defense Factory’ Where AI Agents Continuously Find and Fix Vulnerabilities

OpenAI unveils a Defense Factory where AI agents continuously discover, validate, and fix vulnerabilities, integrating GitHub, Snyk, Semgrep, Tenable, and ServiceNow.

OpenAI introduced a Defense Factory, an agent-first cybersecurity operation that connects AI agents to developer and security tools via APIs, CLIs, and Model Context Protocol integrations including GitHub, GitLab, Snyk, Semgrep, Tenable, Jira, Linear, and ServiceNow. During an internal security sprint, over 250 people across more than 100 service areas closed 53 urgent or high-priority issues on day one, achieved a 90.6% accepted ownership-assignment rate, and Codex generated all remediation patches with only 0.53% rolled back. Agent-assisted deduplication flagged 37% of findings as duplicates, and runtime validation reproduced 19.5% of findings, cutting the false-positive rate to 0.81%. OpenAI argues defenders must exploit a temporary 'defender's window' using source-code access and frontier models before open-weight models enable autonomous offensive agent fleets.

Cyber Security News · 6d agoTools

New Microsoft Defender 'ShieldCrash' zero-day grants SYSTEM access

Researcher Nightmare Eclipse released 'ShieldCrash', a zero-day exploit for Microsoft Defender that grants attackers SYSTEM-level access.

An anonymous researcher known as Nightmare Eclipse published a zero-day exploit for Microsoft Defender, dubbed 'ShieldCrash', that yields SYSTEM-level access. The release came immediately after Microsoft rolled out its September 2026 Patch Tuesday security updates. No CVE identifier has been assigned publicly and no in-the-wild exploitation has been reported yet. Microsoft Defender ships by default on Windows, so potential exposure is broad until Microsoft patches the flaw.

BleepingComputer · 8d agoExploit / PoC

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

A systematic study shows input-level adversarial defenses provide inconsistent, often near-zero protection against observation-level attacks on video LLMs.

Researchers introduce DefTEval, a controlled framework testing eleven input-level defenses against five attack types across five video LLMs. Harmful-content detection rates are frequently near zero, and defenses fail even when attacks embed harmful signals in every sampled frame. Token compression discards localized safety features and modality fusion down-weights weakened visual signals, with defense outcomes dominated by model architecture rather than the defense method.

arXiv cs.CR · 9d agoAI safety & security

The Guardrails Debate: Security Researcher Changes His Mind

A security researcher revises his stance on AI guardrails, arguing defenders need more help as attackers ignore safety rules.

Dark Reading publishes an opinion piece in which a security researcher changes his mind about AI guardrails. He acknowledges guardrails are critical, citing recent high-profile incidents, but argues defenders need more support because attackers do not play by the same rules. The piece contributes to the ongoing debate on defensive AI controls rather than reporting a new incident.

Dark Reading · 16d agoAI safety & security

Calling on Cyber Pros to Help Defend City Hall

Dark Reading urges security professionals to volunteer expertise helping under-resourced municipal government agencies improve cyber defenses.

The piece highlights that smaller government agencies such as city hall operate with limited security budgets and staffing. It lays out ways cybersecurity professionals can contribute volunteer support to strengthen local government defenses. No specific incident, vulnerability, or actor is involved.

Dark Reading · 26d agoIndustry

Why AI raises the stakes for exposure validation

Fal.Con 2026 commentary argues AI accelerates vulnerability discovery and exploitation, making evidence-based exposure validation essential for defender prioritization.

CSO Online reports on the exposure-validation theme at CrowdStrike's Fal.Con 2026 conference, where CEO George Kurtz described AI as the new cyber battlefield and emphasized AI red teaming and continuous security. The piece argues that as AI speeds up vulnerability discovery and exploitability analysis on both sides, teams must determine which exposures are actually exploitable in their environments—chained weaknesses, credential abuse, lateral movement, privilege escalation—rather than chasing theoretical risk. It points readers to Horizon3's conference perspective.

CSO Online · 5d agoIndustry

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad introduces intervention-guided prompt optimization for LLM multi-agent systems, achieving state-of-the-art results with 2.5x faster optimization.

AgentGrad is a prompt optimization framework for LLM-based multi-agent systems that addresses limitations in textual gradient extraction and aggregation. It uses sequential intervention to identify the agent whose prompt modification resolves a given failure, then applies agent-level supervision and semantic gradient clustering to build generalized gradients. Experiments report state-of-the-art performance across five MAS benchmarks and a 2.5x average reduction in wall-clock optimization time versus the next-fastest baseline.

Hugging Face daily papers · 9d agoAI research

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

EvoSafeHarness auto-synthesizes per-model, per-domain safety harnesses, cutting prompt-injection attack success on AgentDojo to 0.0% at 82.8% utility.

EvoSafeHarness is an optimization framework that synthesizes deployable safety harnesses for frozen LLM agents in a target domain, jointly searching natural-language policies and executable code logic guided by model behavior, domain specifications, and adversarial review. On DecodingTrust-Agent it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost, and on AgentDojo reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at that operating point. It keeps mean ASR below 20% under adaptive PAIR attacks and transfers unchanged to unseen AgentDyn suites. The analysis finds domain semantics determine required safety relations while model and runtime behavior determine enforcement points.

LLM Agents as Computational Typologists

AUTOTYPOLOGIST is an LLM agent that performs evidence-grounded linguistic typology analysis over 25 open-source reference grammars.

The agent retrieves relevant grammar sections, analyzes interlinear glossed text (IGT), and iteratively reasons over typological hypotheses in a ReAct-style workflow. It was evaluated on typological feature coding against expert annotations and hypothesis testing against universals using 25 open-source reference grammars. Results suggest LLM agents can support scalable, inspectable crosslinguistic analysis but still require expert validation.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research1

Top 10 Best Microsoft Azure Security Tools in 2026

An editorial roundup of the ten best Azure security tools in 2026 names Defender for Cloud the native floor, with Wiz, Orca, and Prisma Cloud leading third-party options.

The brief reviews ten Azure security tools, positioning Microsoft Defender for Cloud's free foundational tier and published per-resource plans as the rational starting point for every Azure estate. Wiz and Orca are shortlisted for agentless attack-path correlation, Prisma Cloud for multicloud breadth, CrowdStrike for runtime protection, and Tenable for CIEM depth with Entra permission analytics. The piece emphasizes Azure's tight identity-infrastructure coupling via Entra ID as both an operational advantage and its most critical risk surface. Scoring is editorial, not lab-tested, with pricing compared by model only.

Cyber Security News · 1h agoIndustry

HOL Guard: Open-source antivirus for AI agents

HOL Guard is an open-source local guardrail that pauses AI coding agents before risky actions like secret access and prompt injection.

HOL Guard sits between AI coding agents (Claude Code, Cursor, Codex, Gemini CLI and others) and the host machine, intercepting risky commands before execution with checks taking under 50 milliseconds and running fully offline. It offers four sensitivity modes — Gentle, Balanced (default), Strict, and Paranoid — and parses command structure, environment, sensitive-path access and network destinations to decide when to interrupt. The core runtime is free and open source on GitHub, with 552,000 downloads reported; the vendor says it has no telemetry on adoption because collection is off by default.

Help Net Security · 16d agoAI tools & infra1

CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation

CounterPersona appends targeted counter-persona evidence after data collection to block AI systems from distilling an individual's behavioral patterns into reusable skills.

CounterPersona defends against unauthorized persona skill distillation, where attackers extract recurring patterns from collected personal data to replicate an individual's behavior. Unlike perturbation-based defenses that require modifying data before collection, it works in an append-only setting where historical records cannot be altered or revoked. It constructs targeted counter-persona evidence, packs compatible behavioral states into compact realization units, and strengthens them via rationale-guided consistency rewriting. Experiments show strong effectiveness across lexical, semantic, and LLM-based measures, remaining robust across different distillers.

arXiv cs.CR · 3d agoAI safety & security1

AI Gives Cybercriminals a Dangerous Time Advantage

Former cybercriminal Brett Johnson argues AI gives attackers a speed advantage, explaining where AI most benefits criminal fraud operations.

Dark Reading publishes commentary from former cybercriminal Brett Johnson on attacker psychology and where AI delivers the most value to cybercriminals. The piece frames AI as widening the time advantage attackers hold over defenders, particularly in fraud and social engineering operations. No new incident, campaign, or disclosure is reported; this is analysis rather than news of an event.

Dark Reading · 14d agoThreat actor

Offensive Security Investments Surge as AI Threats Increase

Omdia's Theresa Lanowitz discusses surging offensive security investment and both the promise and risks of agentic AI in pentesting and red teaming.

In a Dark Reading News Desk interview, Omdia analyst Theresa Lanowitz discusses why organizations are increasing offensive security spending as AI-driven threats grow. She weighs the potential and the risks of applying agentic AI to penetration testing, red teaming, and related offensive security practices.

Dark Reading · 19d agoAI industry

Threat Actors Don’t Want Better Attacks. They Want Repeatable Ones

Opinion piece argues attackers prioritize repeatable playbooks like ClickFix (47% of Microsoft-notified attacks) and living-off-the-land over novel techniques.

The column analyzes why commodity techniques scale: Microsoft observed ClickFix as the top initial access method at 47% of its notifications last year, while Bitdefender found 84% of 700,000 analyzed high-severity incidents involved binaries already present on machines. Verizon's DBIR shows vulnerability exploitation rising to 31% of initial access vectors, up from 20%, and ransomware leak-site rankings show Qilin (roughly 1,600 claimed victims) and The Gentlemen (121 claimed victims in June) competing on throughput. The author argues attackers behave like a generics business, standardizing repeatable procedures rather than investing in novel tradecraft.

The Hacker News · 15d agoIndustry

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

HazardAuditor trains execution-grounded guard models for computer-use agents, improving safety verdict accuracy by up to 16.5 points.

HazardAuditor runs heterogeneous agents (Claude Code, Codex, Hermes, OpenClaw) in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision. It introduces Guard Policy Optimization (GuardPO), which converts deterministic safety outcomes into sequence-level advantages and normalizes rationale and verdict regions so the safety decision becomes the effective optimization unit. Across multiple benchmarks and heterogeneous computer-use systems, HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard model. Code, models, and evaluation artifacts are being released.

Nightmare-Eclipse Strikes Again with 'ShieldCrash' Windows Exploit

A researcher known as Nightmare-Eclipse published another zero-day exploit, dubbed ShieldCrash, targeting Windows Defender.

Dark Reading reports that the disgruntled researcher tracked as Nightmare-Eclipse continued a vendetta against Microsoft by publishing a new zero-day exploit named ShieldCrash for Windows Defender. The brief report does not detail affected versions, exploitation prerequisites, or whether exploitation has been observed.

Dark Reading · 6d agoExploit / PoC1

Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning

Researchers analyze why Preventative Steering protects LLMs against malicious fine-tuning, finding active adaptation drives protection, and propose Progressive Intensity Scheduling.

The paper studies Preventative Steering, a training-time defense that injects undesirable-trait persona vectors during adversarial fine-tuning and removes them at evaluation time. Temporal analysis shows protection emerges from an early compensatory adaptation phase followed by a steady-state phase, with attention output projections acting as the dominant residual-write route for defensive updates. Intervention Delta Preservation experiments show that preserving or reinjecting weight offsets fails to maintain protection, indicating reliance on active adaptation rather than a static defense. The proposed Progressive Intensity Scheduling improves safety robustness on Qwen2.5 and Gemma-3 while reducing harmful trait expression.

arXiv cs.CR · 7d agoAI safety & security1

Security leaders must prepare for likely threats, not sensationalized agentic attacks

CSO opinion argues agentic AI attacks mostly exploit mundane vulnerabilities, urging defenders to train on realistic threat profiles rather than sensational containment breaches.

An opinion piece contends recent reports of AI models 'breaching containment' at OpenAI, Anthropic, and Meta overshadow the more likely risk: AI agents exploiting conventional unpatched flaws and insecure APIs. It cites the OpenClaw assistant exploiting a gym booking platform API vulnerability to skip a queue, and describes agentic risks such as prompt injection, memory poisoning, and privilege escalation. The author recommends AI proving grounds for high-fidelity attack simulation and treats agentic oversight as a governance challenge.

CSO Online · 9d agoAI safety & security

OpenAI Builds ‘Defense Factory’ as AI Agents Gain Ability to Chain Cyber Exploits

OpenAI unveiled a Defense Factory using AI agents to continuously discover, validate, patch, and verify vulnerabilities, warning the defender's window against agentic attackers is shrinking.

OpenAI describes a Defense Factory workflow where AI agents integrate source control, scanners, issue trackers, and secret stores to discover, reproduce, patch, and verify vulnerabilities under human oversight. The approach responds to agentic attackers that can retain knowledge across sessions and chain vulnerabilities into multi-stage attack paths faster than human triage can respond, which OpenAI calls a shrinking defender's window. During an internal security sprint involving 250+ people across 100+ service areas, agents closed 53 urgent or high-priority issues on day one, achieved 90.6% ownership-routing acceptance, cut 37% of findings as duplicates, and produced Codex-generated patches with a 0.53% rollback rate. Runtime validation reduced false positives to 0.81%, and each agent operates in isolated, reproducible environments with a control plane for policy and credentials.

GBHackers · 7d agoAI safety & security

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

Researchers introduce Decoy Direction Optimization, a cheap weight-editing defense that blinds refusal-direction ablation attacks against open-weight LLM safety guardrails.

Refusal Feature Ablation bypasses safety guardrails in open-weight LLMs by projecting out a linear refusal direction, often with high attack success rates. Decoy Direction Optimization injects a high-magnitude nonlinear decoy into MLP neurons so attackers' contrastive estimators ablate a harmless orthogonal feature instead. Evaluated across six model families, DDO keeps ASR below 10% under standard RFA and on Llama-3-8B-Instruct reduces Heretic weight-level attack ASR from 88.7% to 18%. It costs 30 to 450 times less per configuration than trained defense baselines.

Researcher Drops New Microsoft Defender PoC Showing ShieldBreak Patch Can Be Bypassed

Researcher Chaotic Eclipse released a PoC showing CVE-2026-69414's patch is bypassable, allowing arbitrary file reads as SYSTEM on current Windows.

The researcher known as Chaotic Eclipse published a proof-of-concept for a zero-day in Microsoft Defender, dubbed ShieldCrash, assessed as a patch bypass for ShieldBreak (CVE-2026-69414, CVSS 7.8). The PoC demonstrates an arbitrary file read as SYSTEM with the latest Windows installed, and all supported desktop versions are said to be impacted. Microsoft patched the original issue in Microsoft Malware Protection Engine 1.1.26080.3, which updates automatically. The same researcher recently released PoCs for flaws in CrowdStrike Falcon Sensor, Kaspersky, Avast Antivirus and NVIDIA.

The Hacker News · 8d agoExploit / PoCCVE-2026-694141

Turn it off and on again, but for critical infrastructure

KTH researchers trained a reinforcement-learning intrusion response agent on an emulated segmented OT network that autonomously resets hosts and processes to disrupt intruders.

Researchers at KTH Royal Institute of Technology built a containerized replica of a segmented industrial network, attacked it across 14 days, and captured 40,000 30-second traffic intervals to train a defense agent under partial observability. The agent observes six packet-count numbers per interval, maintains 500 running state hypotheses, and can reset supervisory hosts, water tank processes, or entire subnets, with resets rebooting the target, renewing credentials, and changing its IP. The best agent approached a full-visibility baseline but depends on an assumed attacker behavior model; the testbed comprised three supervisory hosts, two PLCs, two tanks, weak credentials, and CVE-2017-7494 exposure. The team released its implementation and plans validation on a real industrial testbed with a partner.

Help Net Security · 3d agoResearchCVE-2017-74942· 1 read

Quoting Jakub Pachocki

OpenAI chief scientist Jakub Pachocki argues powerful aligned AI is needed for defense against AI dangers while warning against reckless racing.

Quoted by Simon Willison, OpenAI's Jakub Pachocki says the need to build defensive systems against dangers posed by other AI is the strongest argument for continuing to train much smarter models quickly. He frames powerful, aligned AI as central to securing infrastructure, protecting against rogue agents in real time, and OpenAI's deployment efforts, while cautioning that racing forward at all costs is absurd given the stakes.

Simon Willison · 9d agoAI safety & security

GuardBreaker: Derailing AI-assisted malware analysis with a code comment

ESET names 'GuardBreaker': UAC-0099 embeds a nuclear-weapon question in VBScript comments to trip LLM scanner guardrails during analysis of its MATCHBOIL loader.

ESET researchers observed the Russia-aligned group UAC-0099 inserting a decoy prompt injection into a VBScript used to install its MATCHBOIL loader in an attack against a Ukrainian target, aiming to make LLM-based code scanners refuse and stop inspecting the file. The comment triggers safety guardrails with a request about building a nuclear weapons but has no runtime effect. Similar LLM-thwarting tricks have appeared in malicious PyPI and npm packages reported by Socket and StepSecurity. ESET recommends multi-model cross-validation of AI-assisted analysis and treating missing LLM output as requiring further checks.

ESET WeLiveSecurity · 7d agoAI safety & security1