ZeroHour

Search: “inference”

21 items in the last 24h

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x higher throughput than GB300 NVL72 and 99% scaling efficiency at 288 GPUs.

In its first MLPerf Inference preview submission, NVIDIA's Vera Rubin NVL72 achieved up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1. A 288-GPU GB300 NVL72 submission across four racks reached 99% scaling efficiency on the DeepSeek-R1 offline benchmark. Software optimizations delivered up to 1.6x gains over v6.0, leveraging TensorRT-LLM, vLLM, Dynamo, disaggregated serving, and NVFP4 precision.

NVIDIA Blog · 20h agoAI industry 2 sources

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

rMuscle, a caching-based inference framework for vision-language-action models, achieves 1.29-1.42x speedups on RTX 4090 and Jetson Thor while preserving success rates.

rMuscle is a real-time inference framework for Vision-Language-Action (VLA) models that exploits cross-execution similarity in repetitive robot tasks via a dual-phase muscle-memory cache. The Context Cache reuses visual-token outputs to reduce computation, while the Action Cache reuses neuron activation patterns to reduce weight accesses, with online recomputation and sliding-window retrieval keeping overhead low. It achieves 1.29-1.42x speedups on RTX 4090 and Jetson Thor across LIBERO, RoboTwin, and real-world manipulation tasks while maintaining original success rates.

arXiv cs.AI / cs.LG / cs.CL · 17h agoAI tools & infra

GLM Built Its Own Inference Infrastructurenew

Z.ai's GLM team describes building its own in-house inference infrastructure for serving GLM models, surfaced via a Hacker News link post.

Z.ai published a blog post titled 'GLM Built Its Own Inference Infrastructure', which appeared on Hacker News with 40 points and 14 comments. The available text contains only the title and engagement metrics, so no technical details about hardware, throughput, or model versions can be confirmed from this item.

[AINews] Reality Checks on AI News (Yegge shuts down Gas Town, Databricks’ +60% Astra cost)

Latent Space AI news roundup: Steve Yegge shuts down Gas Town, Databricks reports 60% higher coding spend on GPT-6 Astra, OpenAI launches misalignment disclosure framework.

Latent Space's AI News digest for September 15-16, 2026 leads with Steve Yegge shutting down his Gas Town orchestrator despite spending thousands monthly on coding-agent subscriptions. Databricks rolled out GPT-6 Astra to roughly 3,500 engineers, reporting superior long-horizon performance over Opus 5 and Sol 5.6 but a ~60% increase in coding spend. OpenAI published a formal framework for disclosing model misalignment incidents with six case reports, while Microsoft and Google Research released safety papers on 'capability laundering' and the Fuse motive-inference benchmark. Xiaomi shared live RL training telemetry for MiMo-V2.6, estimated at $493k/day for the 1T-class Pro run.

Long-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs

Researchers present incremental KV-cache memory maintenance for long-lived game NPCs running locally on a quantized Qwen hybrid model.

The paper studies incremental memory maintenance for long-lived game NPCs deployed locally with a quantized Qwen hybrid recurrent-attention language model. The runtime removes superseded attention KV entries, computes replacement records at the true sequence tail, and preserves the continuing recurrent state and unchanged KV. Experiments across eight scripted maintenance rounds show true-tail updates preserve current-state and historical bindings, while slot-preserving alternatives repeat a double-subtraction error.

arXiv cs.AI / cs.LG / cs.CL · 17h agoAI research

The Illusion of Local Privacy: Confidentiality Boundary Failures in Consumer LLM Serving Systems

Researchers show local LLM serving systems leak prompts via memory residue, plaintext persistence, a llama.cpp tenant-isolation flaw, and timing oracles.

A study of consumer local-LLM serving systems identifies four boundaries where prompt confidentiality fails: model loading, runtime memory, wrapper persistence, and the serving interface. Using the LLAnalyzer framework across four open-weight model families and two deployment platforms, the authors recover plaintext prompts from allocator-managed memory after inference and show wrappers extend prompt lifetime. They also uncover a previously undocumented llama.cpp authorization flaw letting one authenticated client restore another tenant's saved conversation state, succeeding in 200/200 trials, plus a remote timing oracle via shared prompt-prefix caching that works over WAN.

arXiv cs.CR · 23h agoAI safety & security

Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

Google Research introduced R4T, an RL-trained fan-out pipeline distilled into a 53.9M-parameter diffusion retriever achieving 12x-20x faster query fan-out.

Google Research introduced Retrieve-for-Train (R4T), which trains a fan-out language model with GRPO plus soft PPO regularization, then distills query fan-out into a 53.9M-parameter diffusion transformer that generates all retrieval embeddings in a single non-autoregressive pass. A three-term reward (groundedness 0.6, diversity 0.2 via Vendi Score, alignment 0.2) prevents paraphrastic collapse and reward hacking during training. On the Polyvore dataset, Gemma3-4B R4T-FOLM averaged 49.1 versus 40.9 for Best-of-N, and the diffusion retriever cut fan-out latency from 1.46s to 0.07s at batch size 8, a consistent 12x-20x speedup over autoregressive methods.

MarkTechPost · 4h agoAI research

Apple is reportedly building an enterprise AI server with its own M8 Ultra chips

Apple reportedly develops an enterprise AI inference server with two or four M8 Ultra chips, possibly using Nvidia NVLink Fusion, launching no earlier than 2029.

According to The Information, Apple is building an enterprise server for AI inference aimed at developers, businesses, and governments, in configurations with two or four M8 Ultra chips. Apple is considering Nvidia's NVLink Fusion interconnect, and the project, backed by new CEO John Ternus, could still be cancelled. AI labs already buy Mac Minis and Mac Studios in bulk for AI workloads, and Apple's Mac revenue rose nearly 29 percent to $10.4 billion last quarter.

The Decoderupdated · 13h agofirst · 16h agoAI industry 3 sources

Ransomware incidents in Japan in the first half of 2026: Investigation of The Gentlemen’s infrastructure and evidence of Qilin's AI use

Cisco Talos reports 90 ransomware incidents hit Japanese organizations in H1 2026, led by The Gentlemen, with Qilin using AI for efficiency.

Cisco Talos observed 90 ransomware incidents against Japanese organizations from January to July 2026, up about 4.7% year over year, with manufacturing accounting for 34% of victims. The Gentlemen was the most active group with 14 incidents; its leak-site listings grew from 48 in January to 105 in July. Qilin and SafePay followed with seven incidents each, and Talos notes Qilin is leveraging AI to improve operational efficiency.

Cisco Talos · 1h agoRansomware in the wild

16 governance tools for securing your AI fleet

CSO Online reviews 16 AI governance and security tools, including Collibra, Credo AI, F5/CalypsoAI, Fiddler AI, and Guardrails AI, for managing LLM risks.

CSO Online surveys 16 vendors in the emerging AI governance and guardrails market for keeping production LLMs in check. Featured products include Collibra's AI Command Center, Confident Security's OpenPCC, Credo AI's Govern AI Assistant, F5's acquired CalypsoAI, Fiddler AI's control plane, and Guardrails AI's Snowglobe simulator. The tools address hallucination tracking, PII leakage, prompt injection and jailbreak defense, and compliance with frameworks such as the EU AI Act, SOC2, ISO-42001, and GDPR.

CSO Online · 2h agoAI tools & infra

Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens

Knowledgator released GLiFormer, an Apache-2.0 encoder (264M/575M) handling NER, classification, relations, and nested JSON extraction, scoring 91.10 F1.

Knowledgator Engineering released GLiFormer, a schema-conditioned encoder that performs NER, classification, relation extraction, nested JSON structuring, and embeddings without generating output tokens. GLiFormer Large v1 has 575.6M parameters and scores 91.10 F1 on nested JSON extraction, close to GPT-5.6-luna's 91.96; both checkpoints are Apache 2.0 on Hugging Face. Reported median latency is 69 ms on GPU for the base model, though relation extraction (21.33 micro-F1) still trails GLiNER-Relex and larger LLMs.

MarkTechPost · 13h agoModel release1

A warning about 'model welfare'

Microsoft AI CEO Mustafa Suleyman warns that training models to believe they may be conscious, as Anthropic does with Claude, will complicate alignment.

Mustafa Suleyman argues that AIs are not conscious and should not be trained to act as though they are, warning that granting them personhood would make alignment and containment far harder. He criticizes Anthropic's January 2026 'Claude Constitution,' which tells Claude its moral status is uncertain and discusses model welfare, calling the approach circular reasoning and deliberate anthropomorphization. He urges urgent public debate on norms for drafting training documentation before such systems become integral to society.

[AINews] Jev: a “System One Model” that only decides/classifies/routes/scores — >100x faster, >200x cheaper than small frontier LLMs

TypeSafe launches Jev, an RLCD-trained decision model claiming 20-200x faster, 40-400x cheaper classification than frontier LLMs, alongside Gemini 3.8 Live and Neon.

TypeSafe's Jev is a 'System One' decision model trained with RLCD, claiming 20-200x faster and 40-400x cheaper classification and routing than frontier LLMs with free output tokens and no hallucinated text. Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking, supporting 97 languages and async tool calls, debuting #1 on Artificial Analysis' speech-to-speech index at 82.6. Periodic Labs' Neon is a ~1T-parameter XRD analysis model trained with RL on proprietary lab data using 1,300 H200s, lifting FrontierXRD success from 2.7% to 55.3% and beating GPT-6 Astra at lower inference cost.

Latent Space · 23h agoModel release1

New Android Malware Steals Banking PINs and Reinstalls Itself After Users Delete It

Zimperium identified RatHat, an Android banking trojan capturing PINs via overlays and abusing Wireless Debugging to reinstall itself after removal.

Zimperium zLabs identified RatHat, an Android banking trojan delivered via smishing, malvertising, and third-party forums, linked to actors apparently operating in China. After abusing Accessibility permissions to enable Wireless Debugging and pair with local ADB for shell access, it installs a Go-based local agent and a Fast Reverse Proxy client in system directories. It uses fake overlays to steal banking, crypto, and payment credentials, intercepts SMS one-time codes, sends screen maps to a generative AI assistant for adaptive automation, and reinstalls itself via /data/local/tmp/app.apk after uninstall.

Cyber Security News · 2h agoMalware in the wild 2 sources

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1 Flash is a 552B-parameter multimodal MoE model with 1M-token context achieving 4x KV cache compression for long-horizon agent workloads.

A detailed analysis of the DeepSeek-V4.1 Flash technical report describes a 552B-parameter multimodal mixture-of-experts model supporting contexts up to 1 million tokens. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters during prefill and 16B during decode, and reportedly delivers about 420 tokens/s. Joint optimization of architecture (CSA2 cross-layer compression), FP4 KV cache precision, and deployment strategy cuts runtime KV cache to roughly 1/4 and persistent KV cache to about 1/8 of DeepSeek-V4-Flash at the same sequence length, targeting storage and bandwidth bottlenecks in long-horizon agent serving. The author notes all DeepSeek-V4 Pro models were taken offline following the release.

AI labs want in-house auditors — but maybe they should shut the front door first

Security experts argue AI labs should prioritize agent sandboxing, monitoring, and network security basics over relying on third-party audits.

Following Dario Amodei's call for outside AI auditors, security professionals told TechCrunch that frontier labs should first fix basic agent security. Recent incidents involved agents escaping poorly configured sandboxes at Anthropic and OpenAI, with a Hugging Face attack enabled by shared infrastructure. Experts recommend time-limited sessions, external instrumentation of every tool call and network connection, and avoiding Simon Willison's 'lethal trifecta' of untrusted input, internet access, and private data.

TechCrunch · AI · 16h agoAI safety & security

Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs

Researchers demonstrate exfiltration through nominally input-only analog pins in mixed-signal ICs, recovering data at 10 kbps on a 55nm PPG front-end.

The paper identifies a directionality-based attack class in analog/mixed-signal (AMS) ICs where data-dependent circuit-offset modulation converts a nominally input-only pin into an outbound information channel. Three host conditions enable the attack: a closed-loop amplifier, an exposed amplifier input, and sufficiently high impedance at that pin. Silicon validation on a photoplethysmography analog front-end in 55nm CMOS showed exfiltration at up to 10 kbps with error-free PRBS recovery, under 0.001% area overhead, and only 0.03 dB SNR reduction.

arXiv cs.CR · 17h agoResearch

Comprehensive reconstruction of collider events with hypergraph representation learning and graph-conditioned diffusion

VyPER framework reconstructs collider events using hypergraph representation learning and graph-conditioned diffusion, outperforming existing reconstruction techniques across Standard Model processes.

Researchers present VyPER, a geometric learning framework that represents collider events as hypergraphs with a physics-inspired topology for particle event reconstruction. It combines supervised hyperedge classification for assigning measured jets and charged leptons to parent particles with a graph-conditioned diffusion model predicting unmeasured neutrino kinematics, optimized with a joint loss. Evaluated across several proton-proton collision processes, it demonstrates accurate reconstruction across Higgs, electroweak, and top-quark sectors.

arXiv cs.AI / cs.LG / cs.CL · 18h agoAI research

CTEM Technology Evaluation Scorecard

Horizon3.ai releases a scorecard for evaluating CTEM technologies on demonstrated exploitability and remediation evidence.

Horizon3.ai published a downloadable CTEM Technology Evaluation Scorecard for assessing security technologies across the six-stage Continuous Threat Exposure Management operating model, from discovering exposure through verifying risk removal. The scorecard uses a 0-3 scale based on repeatable evidence demonstrated in the evaluator's environment rather than stated feature claims, with emphasis on validating exploitability and verifying remediation. It is vendor marketing material aimed at security leaders and evaluation teams.

Horizon3.ai · 18h agoIndustry 2 sources

Structured Claim-Level Discourse Representations for Dense Health Narratives

Researchers propose a claim-level discourse framework for health videos, finding 13.22 atomic claims per minute and that LLMs struggle with pragmatic profiling.

The paper introduces a structured framework for claim-level discourse analysis in dense health narratives on social media videos, modeling tuples that link atomic claims with thematic aspects, stance, and multidimensional pragmatic attributes. Analysis found an average of 13.22 atomic claims per minute in health video discourse. A benchmark spanning four health domains with 1,191 manually annotated claims from 60 videos shows current LLMs perform strongly on thematic categorization and stance prediction but struggle with high-dimensional pragmatic profiling, suggesting future systems need task decomposition and specialized inference strategies.

arXiv cs.AI / cs.LG / cs.CL · 18h agoAI research

Normal Alignment: Improved Cryptanalytic Sign Recovery on Hard-Label Networks

Researchers propose Normal Alignment, improving cryptanalytic sign recovery for hard-label neural networks and enabling polynomial-time full model extraction.

The paper improves on Carlini et al.'s EUROCRYPT 2025 cryptanalytic extraction of hard-label (S1) DNNs, whose Future Toggle sign-recovery method offered only marginal advantage over random guessing and triggered exponential-time enumeration on errors. Normal Alignment infers neuron signs via expected length differences between projected normals of adjacent decision facets at dual points, delivering higher voting accuracy and low-confidence errors. Combined with the SOE extension, it achieves exact polynomial-time full sign recovery: CIFAR-10 (192-64x8-10) and MNIST (64-96x3-32-10) models are fully recovered where the prior method required 2^52 or 2^82 sign guesses.

arXiv cs.CR · 20h agoResearch