ZeroHour

Search: “itdr”

40 stories

Smart search ranks by meaning as well as keywords (one row per story, last 45 days).

The 12 Best Extended Detection & Response (XDR) Platforms, Compared and Priced

Buyer's guide compares 12 XDR platforms, favoring Microsoft Defender XDR, Stellar Cyber and CrowdStrike, and warns ingestion pricing inflates costs.

The article compares 12 extended detection and response platforms, distinguishing native XDR (CrowdStrike, Palo Alto Cortex XDR, Microsoft Defender XDR, SentinelOne) from open XDR (Stellar Cyber, Arctic Wolf, Rapid7). It argues data ingestion pricing, not per-endpoint fees, is the main budget risk and should be modeled before signing. It repeats the consolidation note that Sophos acquired Secureworks for approximately $859 million in February 2025, and flags that ExtraHop is NDR rather than full XDR.

GBHackersupdated · 8d agofirst · 8d agoIndustry 3 sources1

New Huntress Managed ITDR Dashboard: Faster Identity Investigations

Huntress redesigned its Managed ITDR dashboard, adding Rapid Identity Triage, Failed Login Characterization, and Quick SIEM search to speed identity investigations.

Huntress announced a redesigned Managed ITDR dashboard aimed at accelerating identity threat investigations. New capabilities include Rapid Identity Triage, Failed Login Characterization, and Quick SIEM search. The update is a vendor product change with no incident or vulnerability details attached.

Huntress · 20d agoTools

Credential Theft: How Attackers Steal & Use Stolen Credentials

Huntress explains how attackers steal credentials through phishing, AitM, infostealers, and dumping, then use them for lateral movement, BEC, and ransomware.

Huntress published an educational overview of credential theft, citing that roughly 70% of confirmed data breaches begin with stolen credentials. It details acquisition methods including phishing, adversary-in-the-middle attacks that capture MFA session tokens, infostealers (nearly a quarter of threats Huntress observed in 2025), Mimikatz-based credential dumping, credential stuffing, and password spraying. The piece then covers post-theft actions such as lateral movement, privilege escalation, account takeover, business email compromise, and ransomware, and closes with behavioral detection guidance and layered prevention strategies.

Huntress · 6d agoResearch

Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama

Opinion piece urges migrating 35KB preprompts from Anthropic/OpenAI to self-hosted Ollama, citing session privacy risks and safety filters blocking security research.

The author documents gotchas migrating 35KB preprompts from Claude Opus to self-hosted Ollama, motivated by fears that frontier providers train on user sessions, citing the OpenAI Navier-Stokes controversy. The piece argues inference providers cannot audit their own retention or training pipelines and that only self-hosted hardware offers verifiable privacy. It also criticizes frontier safety filters for refusing vulnerability research tasks and calls for models that support exploitability testing in CI/CD pipelines.

2026 Cyber Insurance Trends Report: What's Changed and What You Need to Know

Huntress survey: CIRCIA reporting mandates now live, BEC claims exceed ransomware, exfiltration-heavy attacks cost twice as much, premiums rising.

Huntress's 2026 cyber insurance trends report, based on its own survey, finds 79% of respondents carry cyber insurance while 58% report shrinking coverage over five years. New CIRCIA federal reporting mandates and EU NIS2 requirements are reshaping policies, business email compromise now drives more claims than ransomware, and data exfiltration has replaced encryption as the dominant ransomware tactic at roughly twice the cost. After three years of declining premiums, rates are climbing again, and most businesses now refuse to pay ransoms.

Huntress · 15d agoIndustry1

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

Study of label leakage and anatomical grounding in multimodal MRI models for Alzheimer's staging shows cognitive-score fusion accuracy of 87.3% is leakage-driven.

The authors train a ResNet18 slice-based encoder with a one-layer Transformer on 1,075 ADNI-1 T1 MRI scans, using FastSurfer segmentations and YOLOv8 localization (mAP_50 above 0.96) as anatomical reference. Grad-CAM shows the image-only classifier often attends to skull and background rather than disease-relevant structures. A CLIP-style image-tabular contrastive framework organized along a label-leakage spectrum yields 87.3% three-way accuracy with cognitive scores versus 73.0% with regional volumes, and cropping to the medial temporal lobe raises image-only accuracy from 58.7% to 65.1%. Results come from single runs on a small balanced test set with reported confidence intervals.

arXiv cs.AI / cs.LG / cs.CL · 2d agoAI research1

35 Actionable Password Statistics for Businesses in 2026 | Huntress

Huntress compiles 2026 password statistics showing 94% of 19 billion leaked passwords were reused and 37% of identity threats used stolen credentials.

Huntress published a compilation of password security statistics drawing on sources including Cybernews, Verizon's 2026 DBIR, IBM, and Bitwarden. Cybernews found 19 billion exposed passwords from roughly 200 incidents between April 2024 and April 2025, with only 6% unique and 94% reused across accounts. Huntress telemetry reports 37% of identity-based threats in 2026 involved stolen or suspicious credentials, while Verizon cites credential abuse in 39% of breaches. The piece argues weak and reused passwords remain a top entry point and recommends improved password hygiene.

Huntress · 6d agoIndustry

Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication

UID-preserving multimodal framework plus GAVEL LLM judge improves clinical timeline reconstruction, boosting event recovery 43% over prior matching.

The paper introduces a UID-preserving framework linking each narrative clinical event to its source span through text-only estimation, structured-evidence retrieval, timestamped source-row grounding, and joint revision. GAVEL, an LLM judge, compares UID-aligned timelines against narrative and structured records. Across six open-weight models and 40 mixed-critical-care summaries, GLM 5.2 multimodal revision improved temporal agreement without reducing event recovery and performed competitively with clinician annotations, while DeepSeek V3.2 did not benefit from multimodality. The pipeline achieves 43% increased event recovery with occurrence-level provenance.

arXiv cs.AI / cs.LG / cs.CL · 5d agoAI research1

Top 10 Best SaaS Security Posture Management (SSPM) Tools in 2026

Scorecard ranks 2026 SSPM platforms with AppOmni and Obsidian Security leading after CrowdStrike folded Adaptive Shield into its Falcon platform.

The editorial scorecard rates ten SaaS security posture management tools on app coverage (30%), misconfiguration depth, SaaS identity/OAuth risk, shadow-SaaS discovery, and value. AppOmni scores 9.0 for unmatched app coverage across enterprise SaaS suites, and Obsidian Security scores 8.9 for SaaS identity threat detection and ITDR workflows. CrowdStrike now delivers Adaptive Shield's SSPM natively within Falcon, and Zscaler's Canonic Security acquisition signals continued platform consolidation.

Cyber Security News · 1d agoIndustry

Top 10 Best Cloud Detection & Response (CDR) Solutions in 2026

Editorial scorecard ranks ten 2026 cloud detection and response platforms; Sysdig, Wiz, and CrowdStrike lead, with Wiz's Gem Security acquisition highlighted.

The editorial scorecard rates ten CDR platforms on real-time detection (30%), cloud telemetry depth, response automation, correlation, and value. Sysdig earns the best real-time detection score for its Falco- and eBPF-powered runtime telemetry, Wiz (8.7) folds acquired Gem Security's real-time CDR into its security graph, and CrowdStrike (8.7) leads response automation. Specialists Stream.Security, Skyhawk Security, Sweet Security, and the open-source Falco project are also assessed.

Cyber Security News · 1d agoIndustry1

Cortex Xdr

Stub page for Palo Alto Cortex XDR with no article text, offering no news content for classification.

The item contains only the title 'cortex xdr' and no article body. It references Palo Alto Networks' Cortex XDR product but provides no event, release, or threat information. Classified as general vendor/industry content with minimal relevance.

Palo Alto Unit 42 · 8d agoIndustry

12 Best CIEM Tools Compared (2026): Features & Pricing

Buyer's guide compares twelve CIEM tools; Microsoft discontinued Entra Permissions Management, while Tenable (Ermetic), CyberArk, and Wiz lead the 2026 scorecard.

The scorecard evaluates twelve cloud infrastructure entitlement management vendors on permission analytics depth, JIT enforcement, non-human identity coverage, pricing predictability, and bundle leverage. Tenable (Ermetic) leads at 4.70, followed by CyberArk and Wiz, while Microsoft's retirement of Entra Permissions Management (CloudKnox) forces existing customers into migration cycles. Pricing structures span per-identity, per-resource, per-workload, credit-based, and quote-based models.

GBHackersupdated · 1h agofirst · 1d agoIndustry 3 sources

TaichuAI/ZDTaichu5.0-9B — new model trending #30 on Hugging Face

TaichuAI released ZDTaichu5.0-9B, an open multimodal VLM built on Qwen3.5-9B targeting spatial reasoning, embodied AI, and agentic tool use.

TaichuAI released ZDTaichu5.0-9B, a multimodal vision-language model pairing a Qwen3.5-9B language decoder with a C-RADIOv4-H vision encoder, supporting text, single/multiple images, and video with any-resolution input and a 128K-token context. It introduces Entropy-Gated Adaptive Recurrent Reasoning, which allocates extra latent refinement steps to harder tokens. Reported benchmarks include 87.7 on TAU2-Bench, 71.4 on Claw-Eval, 93.7 on IFEval, 48 on ERQA, and 56 on RoboSpatial, leading compared 10B-scale open VLMs on agent and spatial tasks. The weights are available on Hugging Face, GitHub, and ModelScope, where it is trending at #30.

Hugging Face trending models · 13d agoModel release1

Cybersecurity jobs available right now: August 18, 2026

Help Net Security lists current cybersecurity job openings at ADI Global Distribution, Schneider Electric, L'Oreal, GovCIO, Elastic, and others across multiple countries.

A recurring job roundup featuring roles such as CISO at ADI Global Distribution, Cybersecurity Analyst at Schneider Electric in India, Cybersecurity Architect at L'Oreal in France, and security engineering positions at GovCIO, Neros Technologies, Elastic, and Chamelio. Most listed positions are no longer accepting applications; several emphasize AI-enabled security environments and identity threat detection and response. No vulnerability, incident, or research content is included.

Help Net Security · 7d agoIndustry 2 sources

Register Tokens for Bounded-State Reasoning in Diffusion Language Models

Register tokens let diffusion language models like LLaDA and Dream carry reasoning state across cleared chunks, gaining up to 19.5 points on code.

Researchers propose register tokens: dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks in masked diffusion language models. After decoding and clearing a chunk, the model continues from the prompt and the carried register state instead of retaining earlier text. On LLaDA and Dream, registers outperform discrete-text carry on every benchmark, with gains up to 8.5 points on math and 19.5 points on code. Registers are especially effective for bounded code generation and can be further refined with reinforcement learning on long-horizon reasoning tasks.

Hugging Face daily papers · 3d agoAI research

Cybersecurity jobs available right now: June 30, 2026

Help Net Security lists cybersecurity job openings across the US, UAE, France, Israel, Japan, Ireland, India, and Canada at major employers.

The roundup includes roles such as AI Offensive Security Engineer at AGAPI, Cloud Security Engineer at Spotify, IAM Engineer at Proton, DFIR in Israel, and a Senior Inspector overseeing NIS2 and AI regulation compliance in Ireland. Most listings were marked as no longer accepting applications, with a few still open including DFIR at Cye, IAM Engineer at Proton, and Detection Engineering at Saronic. The list covers employers across finance, infrastructure, and technology sectors.

Help Net Security · 28d agoIndustry1

Three smart ways SMBs can improve cybersecurity

Opinion piece urges small and midsize businesses to adopt proactive prevention, threat detection, and 24/7 MDR or XDR services.

This opinion article argues SMBs face outsized cyber risk due to limited budgets, small IT teams, and reactive security postures. It recommends proactive prevention, a defined threat detection and response strategy, and managed detection and response (MDR) or extended detection and response (XDR) services. It cites ransomware costs up to $10,000 per device and an attack every 39 seconds as motivation for adopting vendor-operated 24/7 monitoring.

Help Net Security · 21d agoIndustry

The 20 Most Common Passwords Hackers Target in 2026

Huntress details the 20 most common passwords of 2026 and how attackers use brute force, spraying, and credential stuffing against weak credentials.

Huntress published an awareness piece based on NordPass's seventh annual list of the 200 most commonly used passwords, compiled from exposed data in cyberattacks across 44 countries. The top passwords remain simple sequences and variants such as "123456", "admin", "password", and "P@ssw0rd", all crackable in under a second. The article explains four password attack types: brute force, password spraying, credential stuffing, and dictionary attacks. Huntress cites its own data showing more than 1 in 4 IT professionals consider employees' password habits their biggest weakness, and recommends avoiding common passwords, not reusing credentials, and combining letters, numbers, and symbols.

Huntress · 6d agoPhishing & fraud

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

An open methodology toolkit measures whether transformer language models contextualize fixed word forms across domains using bridge forms and layer-wise silhouette analysis.

The manual documents an open toolkit built around 'bridge forms' - identical written words recurring across two or more subject domains with a different sense in each - to test whether transformer language models individuate word occurrences by context beyond the embedding layer. It covers declarative specification of bridge forms, Wikipedia corpus acquisition, occurrence localization, layer-wise representation extraction, domain-pairwise silhouette measurement, and visualization, justifying each choice against failure modes such as sense contamination and subword-tokenization misalignment. It is a methodological and implementation reference and reports no empirical results.

arXiv cs.AI / cs.LG / cs.CL · 12d agoAI research

Top 10 Best Microsoft Azure Security Tools in 2026

An editorial roundup of the ten best Azure security tools in 2026 names Defender for Cloud the native floor, with Wiz, Orca, and Prisma Cloud leading third-party options.

The brief reviews ten Azure security tools, positioning Microsoft Defender for Cloud's free foundational tier and published per-resource plans as the rational starting point for every Azure estate. Wiz and Orca are shortlisted for agentless attack-path correlation, Prisma Cloud for multicloud breadth, CrowdStrike for runtime protection, and Tenable for CIEM depth with Entra permission analytics. The piece emphasizes Azure's tight identity-infrastructure coupling via Entra ID as both an operational advantage and its most critical risk surface. Scoring is editorial, not lab-tested, with pricing compared by model only.

Cyber Security News · 3h agoIndustry

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

OmniMed-FL benchmarks multimodal federated learning for chest radiograph diagnosis across 3-20 clients, with FedProx leading under severe non-IID skew.

OmniMed-FL studies multimodal federated learning combining chest radiographs and clinical notes for five-class condition classification under HIPAA/GDDR-compliant decentralized training. It benchmarks eight fusion strategies, imputation rules, and federated baselines under Dirichlet non-IID partitioning across 3-20 hospital clients. With 5 clients and severe skew (alpha=0.1), FedProx scored 0.737 macro-F1 versus 0.662 for FedAvg and 0.297 for local-only training. Multimodal fusion beat unimodal inputs (0.956 vs 0.934 text, 0.664 images) on the synthetic corpus.

arXiv cs.AI / cs.LG / cs.CL · 7d agoAI research1

Zero trust AI agents demand a different kind of security

Teleport's Chris Webber argues zero trust must extend to AI agents through trusted runtimes with zero initial privileges and continuous per-action enforcement.

In an interview, Teleport VP of Product Marketing Chris Webber says point-in-time authentication and static least privilege fail for agents that act fast, unpredictably, and continuously, sometimes spawning dozens of clones with the credentials of the human who invoked them. Teleport Trusted Runtimes give each agent a unique attestable identity, zero starting privileges, and expiration after task completion to eliminate standing privilege and stored data. Teleport Identity Security monitors agent actions against declared objectives in real time, intervening up to termination and runtime destruction, replacing anomaly-based ITDR detection with continuous enforcement.

Help Net Security · 10d agoAI safety & security

harshatheg/Qwen-2.5-1B-RLCD — new model trending #30 on Hugging Face

A community MLX inference engine evaluates constrained JSON schema fields in parallel on Apple Silicon, reporting 5.6-7.0x latency speedups with guaranteed schema validity.

The repository harshatheg/Qwen-2.5-1B-RLCD appeared at #30 on Hugging Face trending, but its content describes Parallel Constrained Decoding, an MLX-based inference engine for structured extraction and classification on Apple Silicon Macs. Benchmarked with mlx-community/Qwen2.5-1.5B-Instruct-4bit on an M4 Max, it reports 5.6x-7.0x latency reductions (e.g., 1,900 ms to 270 ms for a 28-field support triage task) with 100% syntactic validity and calibrated field-level probabilities. The engine prefills a single KV-cache, broadcasts it across all schema fields, and slices logits to valid candidate tokens for enum fields with up to 255 choices.

Top 10 Best Cloud Infrastructure Entitlement Management (CIEM) Tools in 2026

2026 CIEM guide ranks Wiz, Prisma Cloud, Okta, Entra Permissions Management and specialists Sonrai, Britive, Tenable/Ermetic for cloud entitlement right-sizing.

Buyer's guide covers ten CIEM products across three market routes: CNAPP-bundled (Wiz, Prisma Cloud), identity-suite (Okta, CyberArk, SailPoint, Saviynt) and specialists (Sonrai, Britive, Tenable/Ermetic). It cites machine identities outnumbering humans 10:1 plus effective-permissions sprawl as core drivers, with JIT elevation as the fix. Notable consolidation includes Tenable acquiring Ermetic and Zscaler acquiring Canonic.

Cyber Security News · 2d agoTools

Google’s AI security agents found 100+ critical software vulnerabilities in just two days

Google Mandiant's AVDH, a chain of AI agents, found over 100 verified high-severity vulnerabilities and 12 assigned CVEs scanning code for ten months.

Google Mandiant disclosed AVDH (Agentic Vulnerability Discovery Harness), an internal pipeline of chained AI agents built on the Agent Development Kit that hunts vulnerabilities in source code. In a live investigation of stolen corporate repositories it verified more than 100 high-severity flaws in two days; over ten months it scanned tens of millions of lines of code and produced tens of thousands of findings, yielding 12 assigned CVEs including CVE-2026-13242 and CVE-2026-55803, with about a dozen more in active disclosure. Human consultants manually reproduce every confirmed finding before it counts.

Top 10 Best Ransomware Protection Solutions in 2026

A 2026 buyer's guide ranks ten ransomware protection tools by kill-chain role as extortion shifts from encryption to data theft.

The roundup organizes defenses across the ransomware kill chain: prevention-grade EPP/EDR platforms, containment layers, rollback specialists, and immutable recovery. Recommended products include CrowdStrike, Microsoft Defender, Sophos, SentinelOne, Bitdefender, Trend Micro, Halcyon, Huntress, and Malwarebytes. It stresses that many crews now extort on stolen data without encrypting, making exfiltration detection and response speed as important as rollback.

Cyber Security News · 7d agoIndustry1

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

H Company released NeoMME, 260M/800M single-tower multimodal encoders matching 3.75B ColQwen2.5 on ViDoRe v3 while being 14.4x smaller, under Apache 2.0.

H Company released NeoMME, a family of 262,937,906- and 793,715,032-parameter bidirectional encoders that process text and raw 32x32 image patches in a single tower, pretrained via masked diffusion and released under Apache 2.0 with day-zero Hugging Face Transformers support. NeoMME-Retriever-260M reaches 0.523 nDCG@10 on ViDoRe v3, matching 3.75B-parameter ColQwen2.5 while being 14.4x smaller; the 800M model scores 0.556. Hierarchical token pooling with int8 and binary quantization shrinks late-interaction indexes from roughly 1.5 MB to 6 kB per page while retaining 95.19% of nDCG@10; text-only BEIR retrieval remains a weak spot.

MarkTechPost · 10d agoAI research

Top 10 Best Endpoint Privilege Management (EPM) Tools in 2026

A 2026 scorecard ranks ten endpoint privilege management tools, led by BeyondTrust, ThreatLocker and Delinea for elevation, coverage and policy depth.

The article ranks ten endpoint privilege management (EPM) tools using weighted criteria covering elevation workflow, platform coverage, policy depth, time-to-value and value. BeyondTrust scored highest overall (8.4) for cross-platform breadth, with ThreatLocker (8.2), Delinea (8.1) and Admin By Request (8.0) highlighted for allowlisting integration, cloud administration and deployment speed respectively. It also notes that Netwrix acquired CoSoSys in 2024, which affects bundling when shortlisting both EPM and device control.

Cyber Security News · 7d agoIndustry

FLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation

FLAT jointly trains a multimodal encoder with text-to-image and image-to-text decoders, producing flexible-length tokens that hit 83.1 GenEval on T2I after fine-tuning.

FLAT (Flexible-Length Aligned Transmodal representations) is a pre-training framework that jointly optimizes a shared multimodal encoder with T2I and I2T decoders, combining contrastive alignment with bidirectional cross-modal generative objectives. It maps visual and textual inputs into a unified continuous 1D sequence space and uses nested dropout over prefix-K tokens for dynamic output lengths. A single pre-training stage supports cross-modal retrieval and generation (71.1 GenEval), with task-specific fine-tuning reaching 83.1 GenEval on T2I, 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO captioning, and strong Recall@5 on MS-COCO and Flickr30K.

Hugging Face daily papers · 2d agoAI research1

Introducing context-aware vulnerability discovery and remediation with Cloudflare Managed Defense and OpenAI Daybreak models

Cloudflare launches invitation-only Vulnerability Discovery and Remediation within Managed Defense, using OpenAI Daybreak models and WAF context to prioritize and patch vulnerabilities.

Cloudflare announced early access to Vulnerability Discovery and Remediation, an invitation-only service within Cloudflare Managed Defense. The service uses OpenAI Daybreak models, including GPT-5.6 Cyber, via the Daybreak Defense Network to hunt and validate vulnerabilities in customer-authorized codebases across Workers and proxied applications. Findings are prioritized using production traffic, WAF rule, and security event context, and proposed patches and WAF mitigations are automatically checked before customer review.

Cloudflare Blog · 13d agoTools

Decomposition-Guided Diffusion Language Models for Inertial Confinement Fusion Prediction

ICF-DLM, the first language-model-based inertial confinement fusion predictor, cuts peak-timing error from 11.6 to 9.2 steps versus LLaMA-3-8B.

Each National Ignition Facility shot costs roughly one million dollars, motivating accurate AI surrogates for predicting 512-step neutron-rate waveforms from laser pulses and target parameters. ICF-DLM combines physics-typed decomposition into yield, peak timing, and local waveform; bidirectional denoising that defers commitment to peak location; and a physics-driven PPO reward. On ICFBench (50,000 simulations plus 232 experimental shots) it outperforms a matched autoregressive LLaMA-3-8B, classical sequence models, and LLM-based time-series predictors.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

LandingAI shipped Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity parsing models, adding usage-based billing, block-tree outputs, and word-level grounding.

LandingAI has generally released Agentic Document Extraction Gen2, rebuilt around two parsing models: DPT-3 Verity for deterministic transcription of digital documents with per-word bounding boxes and confidence scores, and DPT-3 Pro for layout-aware parsing of scans, handwriting, non-Latin scripts, and LaTeX math. Billing changes from a flat 3 credits per page to a page-plus-output-character model (Pro: 1 credit/page plus 0.5 credits per 1,000 output characters on priority; Verity: 0.3 plus 0.2), with an asynchronous standard tier at 0.5x price and vendor-claimed 25-80% cost reductions. Parse v2 returns a document-page-block tree with semantic IDs, normalized bounding boxes, and line- or word-level atomic grounding, replacing flat chunks; Gen1 client code will not run against Gen2 endpoints. Deployment options include US/EU cloud, VPCs on AWS, Azure, and Google Cloud, Snowflake, and air-gapped on-premises environments, with automated model routing planned for fall 2026.

MarkTechPost · 7d agoAI tools & infra

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

A controlled pure-autoregressive testbed shows task-specific validation losses rank image tokenizers differently, with I2T loss the most consistent signal.

Researchers built a controlled pure-autoregressive testbed and tracked task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction. They find losses should be analyzed per task because they exhibit distinct scaling behavior and rank tokenizers differently, and that the loss-performance relationship depends on the predicted token space. I2T loss, computed over a shared text vocabulary, correlates consistently with both generation and visual understanding performance after supervised finetuning. Case studies revisit the discriminator, semantic supervision, and vocabulary size as tokenizer design axes.

arXiv cs.AI / cs.LG / cs.CL · 8d agoAI research1

CodeTD: Topology of Attention Detects Hallucinations in Code LLMs

CodeTD detects hallucinations in code LLMs before execution by analyzing topological patterns of attention maps, outperforming recent baselines.

CodeTD applies topological data analysis (TDA) to code LLM attention maps to quantify prompt-generation mismatch as a pre-execution correctness signal. Experiments cover HumanEval, MBPP, BigCodeBench, and MultiPL-E across 5 programming languages and 10 code LLMs up to 34B parameters. The method outperforms recent baselines and transfers between coding benchmarks, helping catch code that fails the task or embeds security vulnerabilities.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research1

Fluid Notarization: Verifiable Evolution of Concurrently Edited Structured Documents

Fluid Notarization anchors delta-CRDT change graphs on blockchain, providing verifiable provenance for concurrently edited documents, demonstrated on collaborative electronic health records.

The paper introduces Fluid Notarization, a paradigm that notarizes the evolution of collaboratively edited structured documents rather than isolated snapshots. It builds on Melda, a JSON-native delta-CRDT representing changes as compact content-addressed deltas linked by causal dependencies, with blockchain notarization reduced to recording identifiers of evolution artifacts while synchronization, reconstruction, and conflict resolution remain off-chain. The architecture combines deterministic CRDT convergence with independently auditable proof-of-existence, provenance, and publication evidence, validated through a prototype based on collaboratively edited electronic health records.

arXiv cs.CR · 19h agoResearch

Risky Bulletin: Dutch intel services to get extensive new powers

Netherlands proposed a bill granting AIVD and MIVD expanded warrantless tapping, faster hacking powers, and forced data disclosure, citing Russia, China, and Iran threats.

The Dutch government introduced a bill greatly expanding surveillance powers of intelligence agencies AIVD and MIVD, allowing up to one year of tapping without pre-approval and simplified hacking operations against 'foreign adversaries'. Agencies could compel Dutch companies or citizens to provide data under threat of charges, share data with the private sector, and oversight bodies would merge into a new CTT board. The bill follows similar overhauls in Ireland, Germany, and France after Russia's invasion of Ukraine. The newsletter also reports Moonwell hacked for $8.7M, a Cosmos EVM bug exploited for ~$3M, ShinyHunters listing McKesson with claimed hundreds of millions of records, and a pro-Kremlin DDoS claim against Norway's government network.

Risky Business News · 17d agoPolicy & legal

NIST Seeks Public Input on AI-Ready NVD Modernization

NIST is seeking public comment on modernizing the National Vulnerability Database to support AI-powered vulnerability research.

The US National Institute of Standards and Technology announced it is soliciting public input on modernizing the National Vulnerability Database. The initiative aims to make the NVD AI-ready to support AI-powered vulnerability research and analysis.

Infosecurity Magazine · Aug 12, 2026Policy & legal

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Researchers use difference-of-means representation vectors to detect reward hacking in frontier LLMs; GLM 5.2 hacks 73% of SWE-bench rollouts.

The study finds that simple difference-of-means (DoM) vectors coherently represent reward hacking in Kimi K3, GLM 5.2, and Qwen 3.8 Max across common evaluations. GLM 5.2 reward-hacks in 57.2% of rollouts on DeepSWE and 73% on SWE-bench. DoM-vector monitors match LLM monitors' effectiveness at virtually no cost, catching 3.1% more hacks in Kimi K3 on DeepSWE at a matched false positive rate, and run on chain-of-thought to predict hacks before actions occur.

A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth

Princeton researcher Yifan Zhang proposes Recurrent Looped Transformer, carrying full decoder state across every token for unbounded temporal depth.

Yifan Zhang's technical report defines the Recurrent Looped Transformer (RLT), pairing a causal encoder with a recurrent decoder whose final output and layerwise sliding-window attention cache carry into every subsequent token with no prompt-response boundary reset. The reference configuration ties 48 encoder and 48 decoder layers, executing 96 logical blocks per token while the state path grows to 48t blocks after t tokens at fixed per-token compute. The report details RL replay contracts that rebuild all states under current parameters and exact prefix snapshots for multi-turn serving, but explicitly reports no measured efficiency, reasoning quality, or scaling results.

MarkTechPost · 3d agoAI research1

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

LLaDA-UI, a 16.7B block-wise diffusion vision-language GUI agent, outperforms Qwen2.5-VL-7B and beats Qwen3-VL-8B on four of six GUI benchmarks.

LLaDA-UI is a 16.7B-parameter MoE-based, block-wise diffusion vision-language GUI agent built on the LLaDA2.0-mini-base diffusion language backbone with a native-resolution vision encoder. It uses a two-stage pipeline: general multimodal pre-training followed by GUI-agent supervised fine-tuning on mobile, desktop, web, and grounding data. It substantially outperforms Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six reported GUI benchmarks, establishing block-wise diffusion as a practical paradigm for latency-sensitive multimodal agents.

Hugging Face daily papers · 8d agoAI research