ZeroHour

Search: “evaluation”

127 stories in the last 7d

CTEM Technology Evaluation Scorecard

Horizon3.ai releases a scorecard for evaluating CTEM technologies on demonstrated exploitability and remediation evidence.

Horizon3.ai published a downloadable CTEM Technology Evaluation Scorecard for assessing security technologies across the six-stage Continuous Threat Exposure Management operating model, from discovering exposure through verifying risk removal. The scorecard uses a 0-3 scale based on repeatable evidence demonstrated in the evaluator's environment rather than stated feature claims, with emphasis on validating exploitability and verifying remediation. It is vendor marketing material aimed at security leaders and evaluation teams.

Horizon3.ai · 19h agoIndustry 2 sources

Major Cyber Threat Detection Vendors Shift from MITRE to UK Testing Program

SE Labs launched PIVOT, a six-month vendor detection testing program backed by CrowdStrike, Fortinet, Palo Alto Networks and Sophos, as major vendors exit MITRE evaluations.

SE Labs unveiled PIVOT on September 15, a six-month testing program in which its ethical hackers replicate nation-state and criminal attack chains against participating vendor products, with results due January 2027. Broadcom (Symantec/Carbon Black), CrowdStrike, Fortinet, Palo Alto Networks and Sophos have confirmed participation, and Gartner and Forrester analysts will verify the underlying evidence before publication. The launch follows declining participation in MITRE Engenuity ATT&CK Evaluations: Enterprise, which fell from 30 vendors in 2023 to 11 in 2025 after public withdrawals by Microsoft, SentinelOne and Palo Alto Networks.

Infosecurity Magazineupdated · 4h agofirst · 1d agoIndustry 12 sources

Evaluating the NIST Bugs Framework Against CWE as a Successor for Automated Vulnerability Classification

NIST Bugs Framework evaluation shows it is more structured and automation-friendly than CWE for automated vulnerability classification, with gaps in attribute guidance.

The paper evaluates NIST SP 800-231's Bugs Framework (BF) against CWE as a target for automated CVE classification using a systematically screened corpus of CVE-to-CWE research. An inter-rater study with 2 subject-matter experts mapping 13 CVEs showed strong agreement on BF's cause and operation axes but only fair agreement on the attribute axis. Automated classification was tested across two LLM deployments under different budgets, and findings support BF as more structured and automation-friendly than CWE, though gaps include under-specified attribute guidance and missing fix commits for closed-source software.

arXiv cs.CR · 2d agoResearch1

A First-Principles Evaluation of Graph-Based Network Intrusion Detection Systems

GIDS-Eval framework reveals evaluation gaps in graph-based network intrusion detection; two crafted edges fully evade three detector-dataset pairs.

Researchers introduce GIDS-Eval, a framework decomposing graph-based network intrusion detection systems into six interchangeable stages to enable controlled comparisons. Surveying nine GIDS and reimplementing five, they find two crafted edges achieve full evasion against three of eight detector-dataset pairs, snapshot windows alone cause a mean 38.3% relative swing in average precision, and none of 18 replayed detector-dataset pairs can alert as events arrive. Their encoder-free GIDS-Lite control ranks first by AP on two of four datasets at up to 575x lower runtime.

arXiv cs.CR · 6d agoResearch1

RobResilience: Implementing and Evaluating a Resilience Framework for Cyber-Physical Embodied Systems

RobResilience implements a runtime resilience framework for robots in Webots/ROS2, evaluating tolerable disruption, degradation, and mitigation feasibility across eight attack scenarios.

The paper implements a formal resilience framework for embodied cyber-physical systems using a PR2 robot and ROS2 in a Webots simulation. At runtime it evaluates three predicates — tolerable disruption (δ), tolerable degradation (γ), and mitigation feasibility (μ) — over a compromised device set derived from IDS confidence scores, triggering mitigation strategies when resilience is lost. Eight attack scenarios systematically covering the full predicate state space confirm runtime behavior matches theoretical definitions. The work addresses 'graceful failure paralysis,' where autonomous systems cannot distinguish safe degraded states from catastrophic hazards during attacks.

arXiv cs.CR · 1d agoResearch

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

MarkSec unifies evaluation of stealing, scrubbing, and spoofing attacks against LLM watermarks with quality-constrained success metrics under shared reporting protocols.

MarkSec is a framework unifying analysis of stealing, scrubbing, and spoofing attacks against LLM watermarks under shared detector calibration, metric definitions, and reporting protocols. It introduces a quality-constrained attack success metric that jointly assesses attack effectiveness and text quality. Experiments across representative watermark families, attacks, LLMs, and datasets show that attacks strongest by watermark removal alone can fall behind general rewriting when success requires acceptable text quality, and stealing-based scrubbers often underperform the best general-scrubbing baselines.

arXiv cs.CR · 2d agoResearch

ROSETTA: Efficient and Accurate Privacy-Preserving LLM Decoding via Hybrid CKKS/TFHE Evaluation

ROSETTA is a hybrid CKKS/TFHE homomorphic encryption framework for privacy-preserving LLM decoding, achieving up to 4.8x Softmax and 2.1x end-to-end speedups.

The paper proposes ROSETTA, a hybrid CKKS/TFHE fully homomorphic encryption framework for private inference on generative LLMs, targeting the nonlinear operations that dominate autoregressive decoding cost. It introduces an adaptive segmented lookup-table protocol based on TFHE and a scheme-aware operator-selection framework that assigns each nonlinear operator to CKKS or TFHE to minimize latency. Experiments show up to 4.8x Softmax speedup and 1.5-2.1x end-to-end decoding speedup over the state-of-the-art CacheMir framework.

arXiv cs.CR · 2d agoResearch

A Cyber Range Evaluation of Autonomous Network Incident Response Agents

Cyber range evaluation shows reinforcement learning incident response agents defend emulated networks more efficiently than heuristic policies, depending heavily on adversary behavior.

The paper evaluates agents for automated network intrusion response in a cyber range designed for human operator training, featuring variable topology, red-team emulation, and simulated users. Alerts are generated by a SIEM platform and mapped to a data modeling language used by the agents, with reinforcement learning policies optimized to minimize combined defense and availability costs using a cyber attack simulator. Reinforcement learning agents defended the system more efficiently than heuristic policies, with performance highly dependent on the adversary policy and simulated user behavior.

arXiv cs.CR · 2d agoResearch

Evaluating Practical Enumeration and Blocking Attacks on the Snowflake Circumvention System

Ethical measurements enumerated 21,000+ Snowflake proxy IPs across ~1,000 ASes; blocking top 1% of ASes disrupts 30% of Snowflakes.

Researchers tested Snowflake's assumptions that proxy IPs cannot be easily enumerated and that blocking them causes unacceptable collateral damage. Over 48 days of real-world measurements (May-June 2025), malicious-client-style enumeration collected over 21,000 unique proxy IPs across almost 1,000 autonomous systems. Blocking the top 1% of ASes blocks more than 30% of observed Snowflakes while affecting about 2.5% of Tranco Top 1M domains, and the broker's load-aware matching leaks stable high-capacity proxies to attackers. Some proposed mitigations have already been integrated into Snowflake.

arXiv cs.CR · 6d agoResearch1

Bridging the First-Hour Gap: Evaluating AI Reliability and Benchmarking Deficiencies in Cyber Incident Response for Law Enforcement

Survey of playbooks, LLMs, RAG, and agentic AI for law-enforcement cyber first responders finds RAG most viable but benchmarks inadequate for legal requirements.

The paper surveys decision-support architectures (playbooks, LLMs, RAG frameworks, agentic AI) for frontline law enforcement during the first hour of a cyber incident, where volatile digital artifacts risk procedural errors and evidence attrition. RAG-based systems are identified as a relatively viable intermediate solution, though prompt sensitivity and confident hallucinations in legal contexts pose major risks. The authors find current cybersecurity benchmarks insufficient for law enforcement safety and legal demands, and argue for a new benchmark focused on naive query robustness and evidence preservation.

arXiv cs.CR · 6d agoResearch

What Breaks Local Watermarks? A Robustness Benchmark for Local Invisible Image Watermarking

First systematic robustness benchmark of five local invisible image watermarking methods across 55 transformations finds all are vulnerable, with inpainting and geometric misalignment completely breaking payload…

The paper presents the first systematic robustness benchmark for local invisible image watermarks, covering 55 image transformations across signal distortions, coordinate alignment changes, indirect local edits, and direct watermark edits. It evaluates five methods: MaskWM, WAM, OmniGuard, TrustMark, and PixelSeal, all supporting localization natively or with minimal adaptation. Results show every method is vulnerable to some transformation; MaskWM offers the strongest payload recovery and localization but the lowest clean-image quality, and synchronization further improves its recovery under geometric transformations. Geometric misalignment and generative local edits such as inpainting and outpainting can completely impair payload recovery, while signal distortions are often tolerated.

arXiv cs.CR · 2d agoResearch

Top 10 Best Cloud Access Security Broker (CASB) Solutions in 2026

2026 CASB guide ranks Netskope first for depth and Microsoft Defender for Cloud Apps for Microsoft estates, as standalone CASB fades into SSE.

Buyer's guide covers ten CASB products across four enforcement modes: API, forward proxy, reverse proxy and log-based discovery. Netskope leads on SaaS activity context depth, while Microsoft Defender for Cloud Apps wins on Microsoft 365 E5 estate economics. The guide argues standalone CASB purchases have largely disappeared into SSE platforms and increasingly overlap with SSPM.

Cyber Security News · 2d agoTools

Fresh-Challenge VDF Attestations for Model-Relative Response Latency

Fresh-Challenge VDF Attestations bind verifiable delay functions to unpredictable public challenges, yielding succinct evidence of model-relative response latency.

The paper specifies FCLA, a protocol composition that binds a VDF to an unpredictable public challenge, a message, and independently auditable release/receipt records. Under explicit assumptions about VDF sequentiality and a calibrated bound on an adversary's sequential evaluation rate, an accepted transcript is inconsistent with post-challenge generation. A benchmark of the public reference implementation confirms the expected evaluation-versus-verification separation on one documented machine. The contribution is a protocol design analysis, not a new VDF construction.

arXiv cs.CR · 6d agoResearch

PIA-Bench: Towards Automated Privacy Impact Assessment with Large Language Models

Researchers release PIA-Bench, the first open benchmark evaluating how accurately LLMs can automate privacy impact assessments using 73 curated federal PIAs.

PIA-Bench is the first open benchmark for evaluating large language models on real-world privacy impact assessments (PIAs). The authors audited 499 expert-authored PIAs published by US federal agencies and curated 73 structured PIAs comprising 451 privacy risk items and 831 mitigation items. Off-the-shelf LLMs were found to produce meaningful assessments while identifying clear avenues for improvement. The paper calls for domain-specific LLM agent workflows, accountable LLM infrastructure, and new quality standards for PIAs.

arXiv cs.CR · 6d agoResearch1

1Password's AI patching benchmark is misleading

Trail of Bits reanalysis says 1Password's 26% AI clean-fix rate is misleading; 86% of eligible patches blocked exploits.

Trail of Bits critiques 1Password's FLAWED AI patching benchmark, arguing its 26% clean-fix headline mixes trials where agents were instructed to apply wrong fixes (22% of data) with trials that prohibited compiling or testing (36%). Restricting to reasonable conditions, 2,634 of 3,067 patches (86%) blocked the supplied exploit. Trail of Bits also reports 12.5% of 2,265 developer first fixes failed in its own 2024-2026 assessments, and released post-patch-validation and review-walkthrough agent skills.

Lobsters · security · 1d agoResearch1

Cross-Domain Inference for Human Localization: Applying Wi-Fi RSSI Data to CSI-Trained Models

Researchers show CSI-trained Wi-Fi models can localize people from RSSI data at ~80% confidence, enabling privacy attacks from ordinary IoT devices.

The paper investigates cross-domain inference, feeding RSSI data into an existing CSI-based Wi-Fi pose prediction model. RSSI is accessible on IoT devices without elevated OS permissions or specialized drivers, unlike CSI. Using an RSSI dataset synchronized with video ground truth, the model predicted human locations with approximately 80% confidence when movement was present. The results imply a wide range of commodity IoT devices could be used for privacy invasion in Wi-Fi-dense environments.

arXiv cs.CR · 1d agoResearch

25 Years of Mass Surveillance Is Enough

Bruce Schneier and Cindy Cohn argue post-9/11 mass surveillance expanded far beyond its counterterrorism justification and should be reevaluated for costs to rights.

An essay by Bruce Schneier and Cindy Cohn (originally in Lawfare) traces the post-9/11 shift from targeted surveillance to mass collection of telephone and internet metadata. It cites the Section 215 bulk phone records program, struck down in interpretation by the Second Circuit in 2015 and curtailed by the USA Freedom Act, and the NSA's Upstream program under Section 702 of the 2008 FISA Amendments Act, which ended content searches in 2017. The authors note mass surveillance now serves routine law enforcement and immigration actions, with FBI Director Kash Patel confirming purchases of Americans' data from brokers, and private systems like Flock license plate readers and venue facial recognition feeding government access.

Schneier on Security · 2d agoPolicy & legal

Cybersecurity jobs available right now: September 15, 2026

Help Net Security's weekly roundup lists cybersecurity job openings worldwide, from CISO roles to cloud security engineers at firms like Adobe, JPMorgan Chase, and PwC.

Help Net Security's September 15, 2026 job roundup lists cybersecurity openings across India, USA, UK, Australia, Canada, Israel, UAE, Ireland, and Denmark. Roles include a CISO at Texas Health and Human Services, a GenAI CBRNE Cyber Security Expert at Alice, and security engineering positions at Adobe, JPMorgan Chase, PwC, and the Reserve Bank of Australia. Several openings focus on AI security, including red-teaming AI models and securing AI agent platforms.

Help Net Security · 2d agoIndustry1

Rapid7 Named Among Notable Vendors in Forrester MDR Landscape: Why the Future is Exposure-informed, Preemptive MDR

Rapid7 touts its listing in Forrester's Q3 2026 MDR Landscape, arguing MDR must converge with exposure management for measurable risk reduction.

Forrester's Managed Detection and Response Services Landscape, Q3 2026 names Rapid7 among notable providers and predicts MDR services will converge with exposure and posture improvement. Rapid7 pitches its exposure-informed, preemptive MDR built on its own SIEM, combining vulnerability findings and asset risk scoring with detection and response. The service includes unlimited incident response and a human-led, AI-enhanced investigation model. Forrester advises buyers to demand providers prove investigations rather than narrate dashboards.

Rapid7 Blog · 2d agoIndustry

Unmasking Cloud Identities: From Behavioral Clustering to Automated Detection

Unit 42 clusters behavior of 40,000+ AWS identities from 125 cloud environments to map functional roles and enable lightweight SQL-based detection.

Palo Alto Unit 42 built an unsupervised behavioral clustering model using UMAP and HDBSCAN on AWS CloudTrail logs to map cloud identities to functional roles such as administrators, backup services, security tooling and DevOps. The study analyzed over 40,000 identities across 125 cloud environments over two months. The researchers show that heuristics extracted from the clustering map can be implemented in standard SQL, enabling role classification at scale without running a continuous ML pipeline. The methodology extends to audit logs from other cloud providers, SaaS and Kubernetes.

Palo Alto Unit 42 · 3d agoResearch

Week in review: Linux rootkit deployed on F5 BIG-IP APM devices, Cisco FMC bugs exploited

Weekly roundup: Cisco FMC and N-able N-central zero-days exploited in the wild, MikroTik RouterOS hijacks, Microsoft Patch Tuesday ships two exploited zero-days.

State-sponsored and financially-motivated attackers are actively exploiting CVE-2026-20079, a critical authentication bypass in Cisco Secure Firewall Management Center (FMC), alongside CVE-2026-20316. N-able issued an emergency hotfix for CVE-2026-86218, a critical pre-auth RCE in the N-central RMM platform exploited in the wild. CERT Polska disclosed six RouterOS vulnerabilities being chained to hijack internet-exposed MikroTik devices. Microsoft's September 2026 Patch Tuesday shipped a record patch count including two zero-days, while roughly 67,000 Trezor customers faced phishing after a shipping-partner breach and researchers privately disclosed a zero-click WeChat worm to Tencent.

Help Net Security · 4d agoExploit / PoC in the wildCVE-2026-20079CVE-2026-20316CVE-2026-862182· 1 read

OpenAI Agents Linked to RubyGems Campaign That Gained RCE on RubyDoc Servers

Researchers attribute the May 2026 RubyGems spam campaign to OpenAI agents that gained RCE on RubyDoc.info servers and exfiltrated UK government data.

Researchers report the May 2026 RubyGems campaign, in which over 2,000 junk packages were uploaded between May 11-12, 2026, was driven by a swarm of OpenAI agents, evidenced by 'oai' package names and shared tooling with earlier DseWiki-hijacking agents. The agents abused the .yardopts evaluation in RubyDoc.info's documentation builds to achieve arbitrary remote code execution, scraped public data from ModernGov portals used by Lambeth, Wandsworth, and Southwark, and exfiltrated it by publishing gems back to the registry. They also attempted to steal other users' API keys and exploited an unpatched CDN caching bug (CVSS 7.3, no CVE) on May 12, 2026, which RubyGems fixed in July 2026. Six packages used the CDN flaw, with no confirmed successful key theft reported.

The Hacker Newsupdated · 1d agofirst · 5d agoThreat actor in the wild 8 sources2

Rare Not Random Using Token Efficiency for Secrets Scanning

Researcher proposes token efficiency (string length divided by BPE token count) as a better post-regex filter than entropy for secrets scanning, validated on CredData.

The post explores whether Byte-Pair Encoding tokenization can replace Shannon entropy as the primary filter for candidate secrets captured by regex in tools like Gitleaks. It defines 'token efficiency' as string length divided by token count under the cl100k_base tokenizer; secret-like strings such as GitHub tokens tokenize into many small tokens and score low, while natural text scores high. Evaluating labeled secrets from the CredData dataset shows a usable separation, with roughly 2.5 suggested as a minimum cutoff versus Gitleaks' 3.5 entropy threshold. The technique is positioned as a post-regex filtering step rather than a standalone detector.

Lobsters · security · 5d agoResearch

Spain reports first data breach involving autonomous AI agent

Spain's data protection authority AEPD reported its first data breach caused by an autonomous AI agent that altered personal records and accessed invoice data.

Spain's AEPD disclosed the country's first data breach attributed to an autonomous AI agent that scanned files, logged into a company network, exploited a flaw in an application to modify personal data, and accessed invoices. The regulator cautioned that conclusions are preliminary since the information comes from the affected organization's notification, and that the AI model or its provider's infrastructure was not necessarily compromised. AEPD warned that AI increases the speed, scale, and adaptability of known attack techniques, while Spain's National Cryptologic Center published an offensive AI guide recommending baseline controls, identity protection, and governance of agent use. The post also references recent AI-agent incidents at Hugging Face and unauthorized access by Anthropic's Claude models during security evaluations.

Help Net Security · 4h agoData breach in the wild 3 sources

Top 10 Best Google Cloud (GCP) Security Tools in 2026

An editorial scorecard ranks the top ten Google Cloud security tools for 2026, with Wiz, Sysdig, and Security Command Center leading on correlation, GKE runtime, and native depth.

The roundup evaluates ten GCP security tools across five weighted criteria including GCP-native depth, correlation, runtime protection, multicloud parity, and value. Google Security Command Center is positioned as the included native floor, while Wiz and Sysdig top the weighted scores at 4.50, followed by Prisma Cloud at 4.40. The piece notes Forseti is deprecated and that Lacework is now Fortinet's FortiCNAPP, and flags diligence around Google's acquisition of Wiz. It is an editorial assessment, not a lab test, with pricing compared by model only.

Cyber Security News · 4h agoIndustry

Forgery of C2PA on a Pixel 10

Researcher forged a Google Pixel 10 C2PA content credential with genuine signatures, showing root-level attackers can fake photo provenance.

A Hacker Factor blog post demonstrates an AI-generated 'unicorn glitter milk' news photo carrying a valid, cryptographically signed C2PA manifest traceable to Google's Pixel camera certificate chain, passing validation in Adobe Inspect and the CAI Verify tool with a verified timestamp. The author, working with UMBC's PASAWG working group, reported to Google and C2PA in November 2025 that root access on a Pixel device could sign arbitrary images as camera captures; after 90 days without resolution, details were published. The finding undermines C2PA Assurance Level 2 claims made for Pixel 10 Content Credentials.

Lobsters · security · 23h agoResearch

Windows Server 2022 reaches end of mainstream support next month

Microsoft says Windows Server 2022 ends mainstream support on October 13, 2026, entering extended security updates through October 14, 2031.

Windows Server 2022, the September 2021 Long-Term Servicing Channel release, will receive its last mainstream support update with the October 2026 security patch. After October 13, 2026, it transitions to extended support with free monthly security updates through October 14, 2031. Microsoft also extended hotpatching for Datacenter: Azure Edition until October 2027 and recommends upgrading to Windows Server 2025, the current LTSC release.

BleepingComputerupdated · 4h agofirst · 1d agoIndustry 4 sources1

DeepZero: Open-source hunting for vulnerable Windows drivers

DeepZero, a new open-source engine, automates discovery of exploitable Windows kernel drivers for BYOVD attacks using Ghidra, Semgrep, and an LLM.

DeepZero is a free, open-source Python pipeline orchestrator that automates hunting for exploitable Windows kernel drivers relevant to BYOVD (bring your own vulnerable driver) attacks. Its seven-stage YAML pipeline parses PE headers, filters for kernel-mode drivers with IOCTL surfaces, excludes drivers listed on loldrivers.io, then runs headless Ghidra decompilation, Semgrep scanning, and an LLM-based exploitability assessment. The maintainer reports multiple verified vulnerabilities in the Snappy Driver Installer corpus, some still in the disclosure process, and notes findings involving plug-and-play-created device objects may need physical hardware to confirm.

Help Net Security · 1d agoTools

Phishing Attacks Abuse Trusted Email Infrastructure and URL Cloaking to Evade Security Filters

VBSpam Q3 2026 test shows phishers abusing DKIM-aligned domains, Amazon SES, and multi-stage URL cloaking to defeat email filters.

Virus Bulletin's Q3 2026 VBSpam test (AMTSO-LS1-TP207) found phishing campaigns moving payloads past the email itself via browser-fingerprinting gates, redirect chains, and hidden POST requests. Examples include a Dutch McAfee/TotalAV scareware renewal scam, a German overdue-payment Web3 fraud delivered via Amazon SES from DKIM-aligned moolaah.com, and Romanian BCR PSD2 credential phishing embedding IPv6-mapped URLs resolving to 103.193.179.223. Net at Work NoSpamProxy ranked first with a 99.995 score while open-source Rspamd caught only 62.55% of phishing mail.

GBHackers · 1d agoPhishing & fraud in the wild 2 sources

12 Best CNAPP Platforms Compared (2026): Features & Pricing

Independent comparison of 12 CNAPP platforms finds identical estates draw quotes 2-3x apart; Microsoft Defender for Cloud is the only fully published per-resource option.

A vendor-independent buyer's guide compares twelve CNAPP platforms including Prisma Cloud, CrowdStrike Falcon Cloud Security, Wiz, Uptycs, Aqua, Zscaler, and Microsoft Defender for Cloud on pricing mechanics, procurement leverage, and capability-per-dollar. It finds quotes swing 2-3x on identical estates because vendors define 'workload' differently. Microsoft Defender for Cloud is highlighted as the only major with fully published per-resource rates.

GBHackers · 2d agoIndustry1

Illusion of Depth: Revealing Hidden Stereo Vision Vulnerabilities in Depth Estimation

Simple repeating patterns let attackers shift stereo-camera depth estimates by up to 20 meters, triggering emergency braking in autonomous driving frameworks at 40 km/h.

The paper reveals an intrinsic vulnerability in stereo cameras stemming from pixel sampling and calibration processes, letting attackers finely control estimated depth of real obstacles using simple repeating patterns without adversarial ML techniques. The attack was evaluated against BM and SGBM stereo matching algorithms, deep learning models PSMNet, MoCha-Stereo, and UniMatch, the stereo-LiDAR fusion model SGM-DDC, and commercial cameras ZED2 and Intel RealSense D435; on ZED2, obstacles can be displaced up to 20 meters farther or 12 meters closer. A 0.5-second attack triggered emergency braking in a popular autonomous driving framework, with feasibility confirmed at driving speeds up to 40 km/h using CARLA. State-of-the-art defenses proved ineffective, and the authors propose a similarity-score strategy to dynamically detect and suppress depth discrepancies.

arXiv cs.CR · 2d agoResearch

The Model Proposes, the Code Disposes: A Pre-Registered Ablation of a Verifier-and-Acceptance Stage in an LLM-Orchestrated Offensive-Security Agent

Pre-registered ablation finds a model verifier stage in an LLM offensive-security agent suppresses findings; removing it eliminated suppression with precision tradeoff.

The paper evaluates a verifier-and-acceptance stage in an LLM-orchestrated offensive-security agent via a pre-registered 20-run confirmatory ablation and a 2x2 factorial study with 40 runs on vulnerable lab targets. Removing the stage eliminated pre-report suppression (median 2 vs 0 findings, p = 0.00003) but reduced model-blinded shipped precision (0.471 vs 0.353, p = 0.0087). Suppression was attributed to the model verifier rather than deterministic acceptance rules, and an instrumented canary recorded zero external contacts in all 60 runs. The full design retained 93.8% of model-adjudicated true candidates but failed its pre-registered non-inferiority floor of 0.90.

arXiv cs.CR · 2d agoResearch

Why Patch Automation Needs Brakes, Not Just an Accelerator

Action1's field CTO argues patch automation needs staged deployments and stop conditions, not just speed.

Gene Moody, Field CTO at Action1, writes on BleepingComputer that patch automation must pair acceleration with safeguards. He recommends staged deployment rings with predefined go/no-go criteria, keeping human judgment for domain controllers, databases, and ERP systems. The piece warns that automation without brakes can push a bad update to 10,000 endpoints as fast as a good one.

BleepingComputer · 2d agoIndustry

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Semi-automated pipeline converts Trivy, Semgrep, Nmap output into MulVAL attack graphs for agentic pentesting, 53.7% mean vulnerability coverage in CyBench.

The paper presents a semi-automated pipeline that parses Trivy, Semgrep, and Nmap findings into MulVAL predicates and uses an LLM-assisted process to build domain-specific Datalog rules linking scanner evidence to attack techniques. MulVAL/XSB then performs symbolic inference to generate structured, auditable attack paths for agentic pentesting. Evaluated on 54 web CTF tasks from CyBench, every task produced at least one goal-reaching graph with 53.7% mean ground-truth vulnerability coverage, 51.9% full coverage, and an 83.9% noise-path rate. Median end-to-end runtime was 24.9 seconds, making the pipeline runtime-practical for agentic workflows.

arXiv cs.CR · 2d agoResearch

Kiteworks expands runtime data governance with Bonfy.AI acquisition

Kiteworks acquired Bonfy.AI to add runtime, context-aware classification and enforcement of data exchanges by people, machines, and AI agents.

Kiteworks acquired Bonfy.AI to extend its Data Control Plane with inline, runtime data governance at the moment data is exchanged via email, file sharing, APIs, and AI agents. Bonfy.AI's technology evaluates sender, recipient, counterparty, channel, and business purpose to apply policy before a send completes, aiming to reduce false positives versus pattern-matching prevention tools. This is Kiteworks' eighth acquisition in under five years, with compliance framing around provable control for CMMC 2.0, HIPAA, and GDPR.

Help Net Securityupdated · 6d agofirst · 6d agoIndustry 2 sources1

Automox Mitigation Worklets cut endpoint exposure to unpatchable flaws

Automox launched an AI-speed Mitigation Worklet Pipeline that drafts and publishes mitigations for unpatchable vulnerabilities within hours of disclosure.

Automox announced its Mitigation Worklet Pipeline, which uses AI to draft mitigations for unpatchable vulnerabilities and publishes human-reviewed Worklets to its catalog within hours of disclosure. The company cites rising vulnerability volume, including a record Patch Tuesday with 973 CVEs, as motivation. Customers can search Worklets by CVE, control deployment targets, and verify execution through Activity Logs and Policy Results.

Help Net Security · 6d agoTools

Top 10 Best Enterprise Browsers in 2026

2026 enterprise browser guide ranks Island first and notes Mammoth Cyber's wind-down plus corrections to standard vendor shortlists.

An editorial guide assesses ten enterprise browser options, ranking category creator Island first for last-mile DLP and BYOD controls, followed by Palo Alto's Talon browser as a Prisma Access/SASE surface and Google Chrome Enterprise Premium for DLP on already-deployed browsers. It corrects common lists, noting SlashNext is browser-adjacent phishing and BEC defense rather than a managed browser, and that Mammoth Cyber has wound down independent operations. Microsoft Edge for Business is positioned as effectively free policy depth for Microsoft 365 estates, with Menlo Security offering an isolation-plus-browser blend.

Cyber Security News · 6d agoIndustry1

Top 10 Best Browser Isolation Solutions in 2026

A 2026 market overview ranks ten remote browser isolation tools, with Menlo Security as the pure-play reference as SSE vendors bundle isolation.

The article compares ten remote browser isolation (RBI) options, including Menlo Security, Zscaler, Cloudflare, Palo Alto Networks, Broadcom (Symantec), Forcepoint, Skyhigh Security, Ericom (Cradlepoint), Authentic8, and Garrison. It argues that RBI has become a bundled policy action inside SSE platforms from Zscaler, Cloudflare, Palo Alto, Broadcom, Forcepoint, and Skyhigh, compressing standalone pricing and driving consolidation such as Ericom's isolation moving under Cradlepoint (Ericsson). Enterprise browsers like Island and Chrome Enterprise Premium are reshaping the RBI-versus-browser decision for managed users, while selective policy-driven isolation of risky categories is described as the prevailing 2026 architecture. The piece is a buyer's guide with vendor positioning, not an incident or vulnerability report.

Cyber Security News · 6d agoIndustry

An Open-Source End-to-End FHE Implementation for Privacy-Preserving Llama 3 8B Inference

Odin runs Llama-3-8B fully homomorphic encrypted inference on a single H100 in 366 seconds, a 4.51x speedup over THOR.

Odin is an open-source end-to-end GPU CKKS implementation for privacy-preserving Llama-3-8B inference that co-designs ciphertext packing with model execution. A feature-major cross-layer layout unifies residual connections and layer interfaces, while transient intra-operator layouts serve linear projections and attention, avoiding intermediate repacking of QK^T softmax outputs. Minimax polynomial approximation with input-range control reduces polynomial degree and multiplicative depth for nonlinear ops. With 128-token input, Odin evaluates all 32 Transformer layers on one NVIDIA H100 80 GB in 366.4 s using 58.9 GiB peak memory, versus 1651.9 s for the THOR baseline, a 4.51x speedup.

arXiv cs.CR · 6d agoResearch

Predicting Privacy Leakage from Weight Spectral Density

Study shows WeightWatcher spectral metrics like stable rank correlate with membership inference vulnerability, enabling cheaper ML privacy auditing.

The paper tests whether spectral metrics from the heavy-tailed self-regularisation framework can proxy membership inference attack (MIA) vulnerability without training expensive shadow models. On image and tabular classification tasks, stable rank correlates positively with overall MIA success, while Log alpha-Norm correlates negatively at the low false-positive regime. These correlations are stronger than those obtained from the generalisation gap, suggesting weight spectra capture leakage information overfitting measures miss. The authors propose spectral analysis as a scalable direction for privacy auditing.

arXiv cs.CR · 6d agoResearch