ZeroHour

Search: “Censuswide”

32 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

CISOs are feeling the security burden of accelerated AI use

Proofpoint's Voice of the CISO survey finds 85% of CISOs prioritizing AI security, with 80% managing AI risks without proportional resources.

Proofpoint's annual Voice of the CISO report, a Censuswide survey of 1,600 CISOs across 16 countries, found 85% rank securing AI assistants, copilots and automation among top priorities for the next two years. Eight in ten say they must manage AI-related security risk without a proportional increase in resources or expertise, and 78% now consider generative AI a major security risk, up 18% year over year. Board alignment improved to 85%, though nearly 80% still report excessive board pressure; 80% cite human behavior as the biggest cyber vulnerability.

Proofpoint Threat Insight · 7d agoIndustry

Proofpoint 2026 Voice of the CISO Report Finds Cyber Resilience Improving, While AI Expands the CISO Mandate

Proofpoint's 2026 survey of 1,600 CISOs finds improving cyber resilience, rising human risk, and expanding AI responsibilities without added resources.

Proofpoint released its 2026 Voice of the CISO report, a Censuswide-conducted survey of 1,600 CISOs across 16 countries fielded in May 2026. Expected material cyberattacks fell from 76% to 61% year over year and material data loss declined from 66% to 53%, but 79% of CISOs now identify human risk as their biggest vulnerability and GenAI security concerns jumped 18 points to 78%. The report also finds 79% of CISOs expect to manage AI-related risks without proportional resources, and 85% say boards are evaluating cyber risk through a commercial lens.

Proofpoint Threat Insight · 7d agoIndustry

OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution

OracleZoom enables recursive extreme-scale image super-resolution via reference-constrained on-policy distillation, reducing hallucinations at deep zoom scales.

OracleZoom tackles recursive super-resolution, where repeated feeding of predictions back into the same model leaves deeper-scale outputs unsupervised as required source resolution grows geometrically. The framework trains on its own trajectory while carrying the last ground-truth evidence beyond the supervision boundary, combining direct and cross-scale supervision, a no-reference quality objective, a KL-constrained pretrained latent prior, and EMA consistency. Across seven datasets it achieves state-of-the-art SR quality across zoom scales, averaging 0.713 CLIPIQA with larger gains at deeper scales and significantly reduced hallucinations. Code, data, and models are publicly released.

Hugging Face daily papers · 11d agoAI research

Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers

Hugging Face details building and using multi-vector late-interaction embedding models with Sentence Transformers for retrieval workloads.

Hugging Face published a guide on multi-vector, late-interaction embedding models (ColBERT-style) supported through Sentence Transformers. The post covers how practitioners can build and use these models for retrieval and RAG pipelines. It is a developer tooling and technique write-up, not a security advisory.

Hugging Face Blog · 29d agoAI tools & infra1

Hackers Expose Data of 1.2 Million Heights Finance Customers

Heights Finance is notifying over 1.2 million customers that hackers accessed a third-party cloud platform holding contact, bank and government ID data.

Heights Finance, a U.S. consumer lender, discovered unauthorized access on May 7, 2026 to a third-party cloud platform used to store customer data; its internal loan management systems and operations were not affected. Exposed data varies by person and may include contact details, financial and bank account information, government IDs and dates of birth for customers, loan applicants, inquirers, and former borrowers of Curo Management and related brands. The company is offering 24 months of free credit monitoring and identity protection; dark web monitoring found no evidence of publication and no threat actor has claimed responsibility.

Security Affairs · 29d agoData breach in the wild

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Audit of 254 SWE-bench submissions finds top coding-agent entries statistically inseparable, so small leaderboard gaps no longer establish rank.

The paper audits 254 SWE-bench submissions across four splits without running models. On Verified, the top two entries each resolve 396 of 500 instances, and exact paired McNemar tests separate none of the 29 adjacent top-thirty pairs at alpha=0.05. Within-model scaffold score ranges reach 29.8 percentage points, versus an 8.8-point spread among the top thirty. The authors release a five-step audit protocol and recommend reporting comparison-set-specific resolution and model-scaffold provenance.

arXiv cs.AI / cs.LG / cs.CL · 1d agoAI research

Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarch

New work characterizes language generation in the limit via finite witnesses, proves a full separation-width hierarchy, and formalizes all results in Lean.

The paper fully characterizes when language generation in the limit is possible for arbitrary families over a countable universe: each target must admit a finite positive witness such that targets activated by any finite sample share an infinite common intersection. It defines positive separation width and proves every level of the resulting hierarchy occurs, with countable families admitting singleton witnesses and unions of families with infinite common cores requiring unbounded finite witnesses. The characterization, a universal normalization, and a diagonal capture lemma are machine-checked in the Lean proof assistant, with the development maintained on GitHub.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Massive Vietnam-Linked APIS Database Exposes Passport and Flight Data

Exposed Vietnam-linked APIS database held 220.8 million passenger and crew records with passport and flight data from January 2017 to April 2026.

Kinryū Labs discovered an exposed Elasticsearch cluster named 'pax-info' containing 220.8 million passenger and crew records, about 107 GB across 29 indices, hosted on IP space assigned to Viettel in Hanoi. The records include names, birth dates, nationalities, passport numbers, and detailed flight information from airlines across Asia-Pacific, Europe and the Middle East. Researchers reported the exposure on June 3 and the database was secured by June 8 with Singapore Airlines coordinating; no evidence of theft was found, but missing server logs mean access cannot be ruled out.

Security Affairs · 8d agoData breach

Does Syntax Matter? A Graph-Augmented Variational Topic Model for Computational Social Sciences

SCPTM graph-augmented variational topic model shows syntax aids topic diversity and descriptor quality but gains stem mainly from the variational encoder.

The Structural Contextual Probabilistic Topic Model represents corpora as heterogeneous document-word graphs with lexical and syntactic edges processed by a Graph Attention Network inside a VAE for mixed-membership topic distributions. Across four corpora, neural gains in document-topic alignment are attributable to the variational encoder rather than syntax, while graph-augmented variants improve topic diversity everywhere. Dependency paths add value on argumentative deliberative texts but are redundant in technical and institutional registers.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research1

FBI Probes Possible Breach of 153 Million Driver’s Licenses

FBI investigates possible breach exposing up to 170M North American driver's licenses, data sold on Exploit forum via 'Nexus' service, linked to IDScan.net.

The FBI is investigating a potentially massive breach of identity data affecting as many as 170 million North Americans, first reported by Krebs. The 'Nexus' service on the Exploit cybercrime forum claimed to hold over 153 million US and Canadian driver's licenses plus ID cards, travel documents, and medical cards, sourced from an active breach at a major identity verification company. Krebs linked the trove to New Orleans-based IDScan.net, which is investigating. The service went dark shortly after publication; experts warn stolen license data (DOB, address, ID numbers) can't be changed and could enable lifelong identity fraud.

Infosecurity Magazine · 13d agoData breach in the wild

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Study finds zero-shot time-series foundation models underperform on CGM forecasting; fine-tuned Chronos-Bolt cuts RMSE up to 18.4% and dietary context adds signal.

The paper evaluates time-series foundation models for continuous glucose monitoring forecasting across eight public datasets covering Type 1 diabetes, Type 2 diabetes, and non-diabetes populations. Under a unified protocol, zero-shot foundation models did not consistently outperform baselines like Elastic Net and PatchTST, but lightweight fine-tuning did, with fine-tuned Chronos-Bolt reducing RMSE by 6.5%-18.4% in the T1D cohort and 8.6%-18.2% in the non-diabetes/T2D cohort. A residual-based fusion framework adding dietary context from CGMacros reduced overall RMSE by about 3% and postprandial RMSE by about 15% versus CGM-only baselines.

arXiv cs.AI / cs.LG / cs.CL · 6d agoAI research

Loan Depot Data Breach Hits 166

Mortgage lender LoanDepot suffered a data breach reportedly affecting around 166,000 individuals, per the headline.

The headline indicates LoanDepot experienced a data breach impacting approximately 166 (thousand) individuals; the exact figure and scope are truncated. No article text is available, so details on data types exposed or the intrusion method are unavailable.

Infosecurity Magazine · 28d agoData breach

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

An open methodology toolkit measures whether transformer language models contextualize fixed word forms across domains using bridge forms and layer-wise silhouette analysis.

The manual documents an open toolkit built around 'bridge forms' - identical written words recurring across two or more subject domains with a different sense in each - to test whether transformer language models individuate word occurrences by context beyond the embedding layer. It covers declarative specification of bridge forms, Wikipedia corpus acquisition, occurrence localization, layer-wise representation extraction, domain-pairwise silhouette measurement, and visualization, justifying each choice against failure modes such as sense contamination and subword-tokenization misalignment. It is a methodological and implementation reference and reports no empirical results.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

What Does an LLM-Agent Leaderboard Rank Actually Compare?

A methodological study shows close LLM-agent leaderboard rank gaps on SWE-bench and similar benchmarks often do not support superiority claims.

The paper defines an estimand-aware pairwise procedure for comparing agents, checking common support and applying explicit uncertainty rules and practical margins. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are frequently unresolved, and proxy labels or utility rules can change which system is selected. The authors argue a leaderboard score summarizes a released evaluation but does not by itself justify pairwise superiority conclusions.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

Grouped Value Attention stores grouped values and reconstructs content keys via a learned linear map, cutting KV-cache size about 45-47% versus GQA.

GVA stores only grouped values and reconstructs content keys with a learned linear map absorbed into the query at decode time, while a small shared decoupled RoPE channel preserves positional information via a separately cached positional key. At 350M parameters trained on 30B FineWeb-Edu tokens, the 16-dimensional positional variant scores 44.18 average accuracy across five tasks versus 44.36 for GQA and 43.88 for MLA. Custom decoding kernels are in development with an open-source release planned.

Hugging Face daily papers · 9d agoAI research

CTM360 Uncovers Over 3,000 Recruitment Phishing URLs Using Browser-in-the

CTM360's RecruitTrap report documents 3,000+ recruitment phishing URLs using BitB windows to steal credentials and relay MFA.

CTM360 identified over 3,000 phishing URLs impersonating recruiters from more than 50 organizations across 14 sectors in the RecruitTrap campaign. Attacks use Browser-in-the-Browser popups with spoofed address bars to harvest Google and Facebook credentials and relay MFA prompts in real time. About 96% of pages used a Calendly theme, with infrastructure concentrated on AWS EC2 and hidden behind Cloudflare.

The Hacker News · Aug 15, 2026Phishing & fraud in the wild

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution

Researchers propose a hierarchical multi-agent system that cuts ransomware analysis cost by 44% while reaching 96.57% detection accuracy.

An arXiv paper (2609.04820) presents a Cost-Aware Hierarchical Multi-Agent System (HMAS) for adaptive ransomware detection and family attribution. Specialized agents run static analysis first, with dynamic and memory modalities invoked only when confidence is insufficient or specialists disagree; a Meta Orchestrator balances accuracy against computational cost via a cost model, and a locally deployed LLM verifies difficult cases. The system achieved 96.57% accuracy, 0.96 F1-score, and 0.99 ROC-AUC for binary detection, and 0.90 macro-F1 for multiclass family attribution. Average analysis cost dropped 43.97% versus exhaustive analysis, with 56.05% of cases resolved using static evidence alone.

arXiv cs.CR · 12d agoResearch

Driver’s License Data for Sale

Schneier on Security highlights that driver's license data is being sold, underscoring concerns over monetization of driver records and surveillance.

A post on Bruce Schneier's blog is titled 'Driver's License Data for Sale.' The available excerpt contains no article body, so the specifics of the reported data sales are not detailed. The topic concerns the commercial availability of driver's license records, a recurring data-privacy and surveillance theme on the blog.

Schneier on Security · 7d agoData breach

Measuring benchmark optimization in speech recognition

Hugging Face examines how much speech recognition systems overfit benchmarks and how to measure benchmark optimization in ASR.

A Hugging Face post on measuring benchmark optimization in automatic speech recognition, analyzing how model improvements on benchmarks reflect genuine capability gains versus overfitting. It is evaluation methodology research with no direct security impact.

Hugging Face Blog · 26d agoAI research

Welsh environment regulator's FoI blunder exposes diversity data of 2,000 staff

Natural Resources Wales inadvertently exposed diversity data of about 2,000 current and former staff via a 2021 Freedom of Information spreadsheet published online.

Natural Resources Wales confirmed equality monitoring data of roughly 2,000 employees who worked between April 2013 and March 2018 was inadvertently disclosed in a spreadsheet released in 2021 in response to a Freedom of Information Act request. The data may have included ethnicity, disability status, religion or belief, sexual orientation, Welsh language ability, and caring responsibilities — special category data under UK GDPR. The breach was reported to the Information Commissioner's Office, the data was removed and permanently deleted, and NRW says it has found no evidence of misuse. The issue was discovered only after a member of the public alerted the regulator on 23 August 2026.

The Register · Security · 9d agoData breach

UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration

UniH3 unifies hierarchical homogeneity and heterogeneity modeling for all-in-one medical image restoration across modalities and degradation types.

UniH3 introduces a Hierarchical Homogeneity Memory module that distills shared anatomical priors from high-quality images, injected via a Homogeneity-Guided Attention mechanism. A Hierarchical Heterogeneity Balancer mitigates inter- and intra-task conflicts during multi-task optimization. It achieves state-of-the-art on MedIR-2D-500K and MedIR-3D-3D benchmarks for both all-in-one and single-task restoration, with code released on GitHub.

Hugging Face daily papers · 7d agoAI research

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Hugging Face guide fine-tunes a 350M-parameter model with 100 GRPO steps to improve structured output reliability.

A Hugging Face blog post demonstrates fine-tuning a 350M-parameter model using GRPO (Group Relative Policy Optimization) with TRL over 100 training steps. The stated goal is more reliable structured outputs from small language models. No article body was available, so details beyond the title are limited.

Hugging Face Blog · 13d agoAI tools & infra

Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation

A study finds LLM-synthesized CodeQL queries improve average F1-score by 82% over baseline queries, offering scalable vulnerability detection versus direct LLM scanning.

Researchers conducted an empirical study evaluating whether LLMs can synthesize executable CodeQL queries from National Vulnerability Database vulnerability data. LLM-generated queries significantly enhanced baseline CodeQL suites, yielding an 82% improvement in average F1-score across a diverse set of real-world vulnerabilities. A cost-benefit analysis shows direct LLM-based scanning of entire repositories is often computationally and financially prohibitive, while LLM query synthesis offers a scalable and cost-effective alternative for large-scale vulnerability detection.

arXiv cs.CR · 7d agoResearch1

Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers

Hugging Face published a tutorial on training and finetuning multi-vector embedding models using the Sentence Transformers library.

Hugging Face's blog walks through training and finetuning multi-vector embedding models with Sentence Transformers. Multi-vector approaches store multiple vectors per document to support late-interaction retrieval. The post is a practical guide for developers building retrieval pipelines with the library.

Hugging Face Blog · 21d agoAI tools & infra1

220 million traveler records exposed in Vietnam-linked APIS leak

Vietnam-linked APIS Elasticsearch leak exposed 220 million passenger and crew records with passport numbers and flight details spanning 2017 to 2026.

Kinryū Labs discovered an exposed Elasticsearch cluster named 'pax-info' holding 210,318,069 passenger records and 10,465,631 crew records (roughly 107 GB across 29 indices) hosted in Viettel-assigned IP space in Hanoi. The database, apparently operated by a Vietnamese organization, was reachable via a chain of two misconfigurations: a cloud-based path that bypassed an HTTP 401 block and acceptance of default credentials. Exposed data included names, dates of birth, nationalities, passport numbers, and detailed flight information for travelers of many nationalities from January 2017 to April 2026. Access was remediated on June 8 after Kinryū Labs notified Vietnamese authorities, airlines, and CERTs, with Singapore Airlines' security team helping coordinate the response; it remains unknown whether any data was copied by malicious actors.

BleepingComputer · 8d agoData breach1

WarmBloodAban/Minimax-h3_Singularity — new model trending #22 on Hugging Face

Community fine-tune Minimax-h3_Singularity enhances MiniMax-H3 video generation with HDR quality, distant face restoration, and improved motion, trending #22 on Hugging Face.

Minimax-h3_Singularity is a community fusion fine-tune of the MiniMax-H3 multimodal video generation model, built from multiple checkpoints and refined with pruning and weight optimization. It supports Text-to-Video, Image-to-Video, Reference-to-Video, and Video-to-Video workflows in ComfyUI, and claims improvements in HDR clarity, distant face restoration, motion fluidity, and fantasy VFX. The authors recommend pairing it with the minimax_h3_ref2v_turbo_4step_v0.1 LoRA for four-step accelerated inference, and an online demo is available via RunningHub.

Hugging Face trending models · 11d agoModel release7· 1 read

RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments

RegionFed is a gradient-level federated learning framework enabling personalized retail query understanding while matching centralized accuracy with differential privacy.

RegionFed is an architecture-robust federated learning framework for personalized query understanding that operates at the gradient level, using the l2 conflict between regional and global gradients to diagnose heterogeneity and control personalization. Existing parameter-level personalized FL methods collapse on transformers, falling below 10% accuracy on T5, while RegionFed deploys unchanged on T5-Small, T5-3B, RoBERTa, and CNNs. RegionFed-Meta achieves 92.27% across Amazon ESCI, Amazon Reviews, and LEAF-FEMNIST, within 0.23 percentage points of the centralized upper bound, with epsilon-approx-0.60 differential privacy.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research1

Robust Coverless Linguistic Steganography via Sentence Embedding Space with Global Resynchronization

Researchers propose a coverless steganographic framework encoding messages as hierarchical clustering paths in sentence embedding space with a Global Resynchronization Mechanism for robustness.

An arXiv paper proposes encoding secret messages as hierarchical clustering paths in the sentence embedding space rather than token space, improving decoding stability against word- and sentence-level textual perturbations. A Global Resynchronization Mechanism (GRM) reframes variable-length bitstreams as discrete symbols anchored to semantic subspaces to prevent bit-slippage. Experiments show substantial robustness improvements while maintaining embedding capacity and resistance to statistical analysis.

arXiv cs.CR · 12d agoResearch

YuE2 · Frontier Music with Symbolic Planning

YuE2, a 3.59B-parameter music generation model, scores 6.9632 on SongBench, beating Suno v5 via symbolic planning.

YuE2 is a music generation model of roughly 3.59B parameters and 28 layers supporting song creation, covering, and agentic editing through editable ABC symbolic scores. Its best-of-8 setting reaches 6.9632 on SongBench, the highest mean among 15 evaluated settings on WildSongBench (192 prompts), ahead of Suno v5 at 6.8721. The project also introduces MERT2, whose 632M-parameter encoders achieve state of the art on 14 of 15 MARBLE metrics, and SheetSage2, which transcribes beats, downbeats, key, chords, structure, and melody with SOTA on 10 of 13 benchmark metrics.

Hacker News · AIupdated · 5d agofirst · 5d agoModel release 2 sourcesHN 43↑ · 35 comments

Type Diversity Enables Transformers to Generalise Compositionally

Researchers show lexical-versus-structural compositional generalization gaps in Transformers stem from type diversity imbalance in datasets, not architectural limits.

The paper argues that Transformers' difficulty with structural compositional generalization is an artifact of low structural type diversity in prior benchmark datasets rather than an architectural limitation. Using Grammatical Framework, the authors create linguistically diverse variants of COGS and SLOG. They find type diversity correlates with compositional generalization equally in lexical and structural test cases, contradicting previous claims that compound divergence explains task difficulty.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research

Condé Nast Data of 32.8 Million Users Offered for Sale After WIRED Leak

A 32.8 million-record Conde Nast user database is offered for $15,000 on a Russian cybercrime forum, extending December's WIRED leak with millions of unseen records.

A database of 32,815,767 Conde Nast user records went on sale on 7 September 2026 for $15,000 on a Russian-language forum, containing names, addresses, birth dates and phone numbers but no passwords or payment data. Ransomnews verified a 5,000-record sample as genuine account data collected between September and late October 2025, with roughly 30.5 million non-WIRED records never previously published. The listing matches the December 2025 WIRED leak of 2,366,576 records, claimed by an actor called 'Lovely' who said 40+ million records were stolen via IDOR and broken access controls. Conde Nast has not confirmed the breach; exposed data enables credible targeted phishing and fraud.

Security Affairs · 9d agoData breach in the wild