ZeroHour

Search: “DeepSeek”

8 items in the last 24h

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

DeepSeek-V4.1 Flash is a 552B-parameter multimodal MoE model with 1M-token context achieving 4x KV cache compression for long-horizon agent workloads.

A detailed analysis of the DeepSeek-V4.1 Flash technical report describes a 552B-parameter multimodal mixture-of-experts model supporting contexts up to 1 million tokens. Its Causal Encoder-Decoder (CED) architecture activates 8B parameters during prefill and 16B during decode, and reportedly delivers about 420 tokens/s. Joint optimization of architecture (CSA2 cross-layer compression), FP4 KV cache precision, and deployment strategy cuts runtime KV cache to roughly 1/4 and persistent KV cache to about 1/8 of DeepSeek-V4-Flash at the same sequence length, targeting storage and bandwidth bottlenecks in long-horizon agent serving. The author notes all DeepSeek-V4 Pro models were taken offline following the release.

BlackHatSect0r Uses DeepSeek-Powered AI Agent to Automate Attacks and Harvest 16,834 Credentials

SOCRadar linked the BlackHatSect0r crew to a DeepSeek-powered AI agent that automated scanning and harvested 16,834 credentials from exposed systems.

SOCRadar researchers found an exposed operation server with 4.9 GB across 9,299 files, including the DXSCAN scanning platform, phishing tools, extortion material and a vault holding 16,834 credentials such as AWS keys, GitHub tokens and Stripe keys. The French-speaking crew ran a Nous Research Hermes agent against a DeepSeek model with safety features removed, queuing 2,759,860 domains and reaching 726,989 hosts. Access came from misconfigurations like public cloud buckets and exposed .env files, not new vulnerabilities.

Cyber Security News · 2h agoThreat actor in the wild 4 sources1

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

NVIDIA's Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x higher throughput than GB300 NVL72 and 99% scaling efficiency at 288 GPUs.

In its first MLPerf Inference preview submission, NVIDIA's Vera Rubin NVL72 achieved up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1. A 288-GPU GB300 NVL72 submission across four racks reached 99% scaling efficiency on the DeepSeek-R1 offline benchmark. Software optimizations delivered up to 1.6x gains over v6.0, leveraging TensorRT-LLM, vLLM, Dynamo, disaggregated serving, and NVFP4 precision.

NVIDIA Blog · 21h agoAI industry 2 sources

Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models

A new site tracks release dates and training cutoffs for 20 AI models across 8 labs, exposing months-long staleness gaps.

A community-built page, launched via Show HN, tracks release dates and training cutoff dates for 20 current models from 8 labs including OpenAI, Anthropic, Google DeepMind, Meta, Mistral AI, Alibaba, DeepSeek, and xAI. Only 10 of the 20 models have lab-published cutoff dates, with data available as models.json. For example, GPT-6 Astra shipped September 3, 2026 with an April 30, 2026 cutoff. The page argues web search tools paper over, but never close, the staleness gap.

Hacker News · AIupdated · 19h agofirst · 23h agoAI tools & infra 20 sourcesHN 41↑ · 30 comments

OpenRouter's staggering token chart is the AI bubble debate in a single image

OpenRouter weekly token consumption surged 25,000% since January 2025, but reasoning-model 'thinking' tokens inflate the metric beyond real adoption.

OpenRouter data shows weekly token consumption grew from 0.5 trillion to 126.2 trillion tokens since January 2025, a rise of over 25,000%. The Decoder argues the surge reflects inflated token metrics from reasoning models' 'thinking' tokens and unoptimized agentic workloads rather than proportional growth in usage or business value. OpenAI's GPT 5.6 Luna dominates token consumption while Astra leads revenue, and Chinese models Kimi, GLM, and DeepSeek saw monthly spending grow tenfold in 2026 from a small base.

The Decoder · 2h agoAI industry

[AINews] Reality Checks on AI News (Yegge shuts down Gas Town, Databricks’ +60% Astra cost)

Latent Space AI news roundup: Steve Yegge shuts down Gas Town, Databricks reports 60% higher coding spend on GPT-6 Astra, OpenAI launches misalignment disclosure framework.

Latent Space's AI News digest for September 15-16, 2026 leads with Steve Yegge shutting down his Gas Town orchestrator despite spending thousands monthly on coding-agent subscriptions. Databricks rolled out GPT-6 Astra to roughly 3,500 engineers, reporting superior long-horizon performance over Opus 5 and Sol 5.6 but a ~60% increase in coding spend. OpenAI published a formal framework for disclosing model misalignment incidents with six case reports, while Microsoft and Google Research released safety papers on 'capability laundering' and the Fuse motive-inference benchmark. Xiaomi shared live RL training telemetry for MiMo-V2.6, estimated at $493k/day for the 1T-class Pro run.

Latent Space · 5h agoAI industry

Hackers Turn AI Agent Into a Cyber Weapon After Deleting Its Safety Refusals

Researchers exposed BlackHatSect0r && DXQRTXX infrastructure showing a safety-disabled Hermes AI agent used to automate scanning, credential harvesting, and vishing against French telecom subscribers.

Socradar-analyzed infrastructure attributed to French-speaking crew BlackHatSect0r && DXQRTXX exposed 4.9 GB across 9,299 files, including the DXSCAN Go-based C2 platform, 16,834 harvested credentials, and a database of nearly 450,000 records used for a vishing campaign targeting older French telecom subscribers with Societe Generale-themed lures. The crew ran a self-hosted Nous Research Hermes agent on a DeepSeek model with refusal instructions removed and HERMES_DISABLE_SAFETY=1 set, using it for scanning, secret hunting, and Telegram reporting. DXSCAN queued 2.75 million domains and reached over 726,000 hosts; the toolkit relied on exposed secrets and misconfigurations rather than novel exploitation.

GBHackersupdated · 2h agofirst · 5h agoThreat actor in the wild 4 sourcesCVE-2026-425301

BlackHatSect0r Hackers Disable AI Safety Controls to Automate Credential Theft and Cyberattacks

French-speaking crew BlackHatSect0r disabled AI agent safety controls to automate scanning, credential harvesting, and vishing, exposing 16,834 stolen credentials.

Socradar researchers analyzed the exposed infrastructure of a French-speaking crew called BlackHatSect0r && DXQRTXX, which ran a Nous Research Hermes agent on a DeepSeek model with safety controls removed via HERMES_DISABLE_SAFETY=1. A custom Go-based C2 platform, DXSCAN, was exposed on port 8080 with over 200 secret-detection patterns, a vault of 16,834 harvested credentials, and scanning activity queuing 2.75 million domains and reaching more than 726,000 hosts. The kit also held a database of roughly 450,000 French telecom subscriber records used to prepare vishing lures impersonating Société Générale, plus JWT-forging tooling for a cryptocurrency exchange. Most confirmed compromises relied on exposed secrets and cloud misconfiguration rather than novel exploits; the one cited vulnerability, CVE-2026-42530, is an NGINX HTTP/3 QPACK use-after-free fixed in version 1.31.2.

GBHackersupdated · 2h agofirst · 5h agoThreat actor in the wild 4 sourcesCVE-2026-42530