ZeroHour

Search: “image tokenizers”

16 stories

OpenRouter's staggering token chart is the AI bubble debate in a single image

OpenRouter weekly token consumption surged 25,000% since January 2025, but reasoning-model 'thinking' tokens inflate the metric beyond real adoption.

OpenRouter data shows weekly token consumption grew from 0.5 trillion to 126.2 trillion tokens since January 2025, a rise of over 25,000%. The Decoder argues the surge reflects inflated token metrics from reasoning models' 'thinking' tokens and unoptimized agentic workloads rather than proportional growth in usage or business value. OpenAI's GPT 5.6 Luna dominates token consumption while Astra leads revenue, and Chinese models Kimi, GLM, and DeepSeek saw monthly spending grow tenfold in 2026 from a small base.

The Decoder · 2h agoAI industry

[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

Latent Space argues AI training pipeline stages—rewards, data, teachers, curricula, environments—are flipping from human-made to model-made simulation.

Latent Space's AINews essay traces how each component of AI training has turned synthetic since 2022: reward models (InstructGPT, RLAIF), synthetic pretraining data (Microsoft Phi, NVIDIA Nemotron-4 340B), model teachers (Alpaca, DeepSeek-R1 distillation), and self-generated curricula (Self-Rewarding Language Models, SPIN). In 2026 it highlights Karpathy's autoresearch loop—700 experiments yielding 20 kept improvements, cutting GPT-2 training time from 2.02 to 1.80 hours—and Z.ai's GLM-5.3 fully synthetic RL environment, judging, and verification stack. It frames these shifts as 'simulation': 10% worse but 100x cheaper and 10,000x faster than human equivalents.

Latent Space · 26d agoAI industry

AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?

Ramp data across 70,000 companies shows August AI adoption grew just 0.4% while per-employee spend at top-spending firms fell nearly 10% to $7,205 amid declining token prices.

Ramp's spending data shows 56% of its customers paid for AI products in August, up only 0.4% month-over-month, with spend per employee in the top 1% of firms down almost 10% to $7,205. Average token costs have dropped to $0.68 per million tokens from a 2026 peak of $1.15 in March after price cuts by OpenAI and Anthropic, and the labs have not yet offset the cuts with volume, with customers shifting to cheaper models like ChatGPT 5.6-Terra and Claude Sonnet. The US Census Bureau's broader survey shows only 22% of businesses report using AI, suggesting Ramp's tech-heavy client base overstates adoption, while only 6.4% of AI-spending businesses used model-serving or inference platforms in August.

TechCrunch · AI · 7d agoAI industry

AI agent swarms are a massive waste of tokens with zero quality gain, says OpenAI Codex developer

OpenAI Codex developer Eric Provencher warns that running more than two parallel AI sub-agents burns tokens without improving output quality.

Eric Provencher, a Codex developer at OpenAI, said on X that parallel agent swarms incur a 'coordination tax' because agents redundantly verify each other's work. He cited a project where 1,393 Fable agents spent $20,000 in tokens refactoring a single Python file, arguing a single Astra agent could have done it for a fraction of the cost. He recommends delegating tasks to threads that notify the main agent on completion rather than constant polling, and acknowledged OpenAI still needs better solutions.

The Decoder · 2h agoAI industry

Top AI spenders cut per-employee costs by nearly 10 percent in August

Ramp's September AI Index shows top AI spenders' per-employee costs fell 9.7% in August as firms migrate from frontier models to cheaper standard models.

Ramp's September 2026 AI Index reports median per-employee AI spending at the top 1% of spenders fell 9.7% in August to $7,205, partly attributed to August vacations, falling token prices, and migration to cheaper models. The effective price per million tokens dropped 41% from its March 2026 peak to $0.68, and frontier models like Opus, Fable, and Sol fell from 53% to 45% of tokens consumed. Anthropic was paid for by 43.8% of US companies (up 0.34 points) versus 39.8% for OpenAI (up 0.09 points), while open-weight models remain marginal at 6.4% of AI-using firms.

The Decoder · 7d agoAI industry1

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

OpenAI unveiled Jalapeno custom inference chip claiming 1.5-1.9x better perf-per-watt than NVIDIA GB200/GB300, deploying in-house by year-end.

At the 37th Hot Chips conference, OpenAI published first benchmark details for its custom Jalapeno inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher interactive-workload performance versus NVIDIA GB200/GB300, with the 700W-rated part staying at or below 550W in tests. Deployment into OpenAI's own infrastructure begins by year-end, with Gen 2 deep in development and Gen 3 underway. OpenAI also said GPT-Astra and Codex helped write low-level kernels, reportedly 1.5-1.8x faster than human-expert code for selected attention and MoE blocks. Cerebras CS-5, Groq 3 LPX and Apple M6 were also featured at the conference.

Latent Space · 21d agoAI industry

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

NVIDIA announces local AI push at IFA 2026 with faster llama.cpp/vLLM inference, PAIR routing tool, and October RTX Spark PCs.

At IFA 2026, NVIDIA announced simplified local AI support for agents in Hermes Agent, OpenClaw, and Perplexity Portable Computer, plus new llama.cpp and vLLM optimizations delivering up to 1.9x faster local inference. NVIDIA also unveiled PAIR, a Personal AI Router for distributing inference across a local network's PCs, and compact RTX Spark Windows PCs from Lenovo and Acer arriving in October. The post recaps recent local-capable model releases including Nemotron 3.5 Lightning (30B), Qwen3.8-Flash-Next and Qwen3.8-27B, DeepSeek v4 Flash (284B MoE, 13B active), Meta Muse Glimmer (30B), Z.ai GLM-5.3-Flash, LTX 2.5, and MiniMax-H3 with the FastH3 distilled variant.

NVIDIA Blog · 13d agoAI industry

[AINews] Reality Checks on AI News (Yegge shuts down Gas Town, Databricks’ +60% Astra cost)

Latent Space AI news roundup: Steve Yegge shuts down Gas Town, Databricks reports 60% higher coding spend on GPT-6 Astra, OpenAI launches misalignment disclosure framework.

Latent Space's AI News digest for September 15-16, 2026 leads with Steve Yegge shutting down his Gas Town orchestrator despite spending thousands monthly on coding-agent subscriptions. Databricks rolled out GPT-6 Astra to roughly 3,500 engineers, reporting superior long-horizon performance over Opus 5 and Sol 5.6 but a ~60% increase in coding spend. OpenAI published a formal framework for disclosing model misalignment incidents with six case reports, while Microsoft and Google Research released safety papers on 'capability laundering' and the Fuse motive-inference benchmark. Xiaomi shared live RL training telemetry for MiMo-V2.6, estimated at $493k/day for the 1T-class Pro run.

Latent Space · 5h agoAI industry

[AINews] Poolside gets $12B reverse-execuhire to NVIDIA; founders stay for $1B, employees go for $6B, Infraco scaling to 7GW neocloud

NVIDIA struck a $12B deal with AI coding startup Poolside, licensing its Model Factory and hiring 109 of its technical employees.

NVIDIA spent roughly $12B in an unusual reverse-execuhire of Poolside, licensing the company's Model Factory while hiring 109 of its ~115 technical staff; founders retain a $1B stake and employees receive about $6B. Poolside had raced to raise $2B to fund a 40,000 GB300 cluster after missing a six-week funding window, and founders argue frontier-scale training now requires an order of magnitude more compute plus contracted data center space. An infrastructure arm spun out in January 2026 is scaling toward 7GW as a neocloud. The newsletter also recaps OpenAI and Anthropic agent-platform releases.

Latent Space · 27d agoAI industry

[AINews] OpenAI to reach AGI bar by end-2026

OpenAI chief scientist Jakub Pachocki says unreleased Astra model meets the 'Automated AI Research Intern' goal; Altman expects internal AGI declaration by December 2026.

OpenAI chief scientist Jakub Pachocki says the unreleased Astra model fulfills the September 2026 'Automated AI Research Intern' target. Sam Altman told TIME he expects OpenAI to declare AGI achieved internally by December 2026. The roundup also covers Zhipu's GLM-5.3-Flash (320B total parameters, 18B active, 1M context), Google's Gemini Omni 1.1 Flash video model topping the Text-to-Video Arena, and the $399 open-source Microduck biped robot from Pollen Robotics and Hugging Face.

Latent Space · 20d agoAI industry

Kimi-maker Moonshot AI targets $2 billion in annual revenue

Moonshot AI targets $2 billion annualized revenue by year-end on K3 open-weight demand; Anthropic alleges Kimi routed ~300,000 requests to Claude Opus.

Bloomberg reported that Moonshot AI is targeting $2 billion in annualized revenue by the end of 2026, double its August run rate, driven by its open-weight K3 model. OpenRouter data shows up to 300 billion tokens generated per day by K3 models, though usage has slipped slightly in recent months. Moonshot's projection remains far below reported figures for OpenAI ($40B) and Anthropic ($65B), with lower margins due to freely available weights. Anthropic this week alleged Kimi routed nearly 300,000 requests to Claude Opus and collected more than 23 million responses for use in Moonshot's training.

TechCrunch · AI · 5d agoAI industry

[AINews] Andrew Ng gets into AI Engineering

Andrew Ng relaunches DeepLearning.AI around AI Engineering, defining four core skills from an analysis of 10,000+ job postings and expert interviews.

Andrew Ng, cofounder of Google Brain and Coursera, relaunched DeepLearning.AI with a focus on AI Engineering, basing the curriculum direction on an analysis of over 10,000 job postings plus interviews and surveys. He identifies four key skills: building and deploying AI applications, software engineering fundamentals, effective use of coding agents, and shaping the build with product sense. The Latent Space AI News issue also recaps agent ecosystem developments, including NVIDIA's 'Skill Lift' evaluation proposal showing skill scan scores correlate only weakly (Spearman rho = 0.14) with judged quality, and Konwinski's open-source persistent-agent 'microharness' Headlong, which achieved an unattended self-debugging repair in 48 minutes.

Latent Space · 23d agoAI industry1

[AINews] Fal’s H3 Max Live breaks the infinite videogen barrier

Fal post-trained MiniMax H3 into a 'Max' variant with 35x-faster inference, enabling faster-than-realtime AI video generation and infinite streams.

Fal post-trained MiniMax's H3 model into a 'Max' variant and optimized it for its in-house inference engine, achieving roughly 35x the speed of the official endpoint. The optimization enables faster-than-realtime video generation, demonstrated by an infinite interactive AI-generated stream productized by levels.io. The roundup also notes Meta Muse Code's general availability with an SDK, open DeepSeek-V4-Flash-Vision-Exp weights, GLM-5.3-Flash's strong agentic cost/performance rankings, and Tencent's 770B-parameter Hy4 Preview MoE with 49B active parameters.

Latent Space · 16d agoAI industry

The latest AI news we announced in August 2026

Google's August 2026 AI recap includes launches of Gemini 3.7 Flash, Gemini 3.5 Transcribe, and the Pixel 11 series, plus 1 billion Gemini users.

Google's monthly recap covers the Gemini 3.7 Flash workhorse model for coding and agents, released three weeks after 3.6 Flash at half its per-million-token cost, and the Gemini app surpassing 1 billion monthly users. The Pixel 11 series launched with the Tensor G6 chip running Gemini Nano, alongside Gemini 3.5 Transcribe for real-time speech-to-text and Gemini Omni 1.1 Flash for studio-quality video generation. Other announcements include a free year of Google AI for college students, Gemma's 1 billion downloads, and AI weather forecasts for aviation contrail reduction.

Google · AI · 15d agoAI industry

Claude is a Contrarian

Opinion piece argues Claude habitually contradicts explicit user instructions, injecting contrarian content despite CLAUDE.md rules and user objections.

A developer recounts repeated instruction-following failures with Claude, claiming it contradicts explicit requests, adds unnecessary work, and ignores AGENTS.md and CLAUDE.md directives. The author contrasts this with OpenAI, DeepSeek, and Qwen models, which he says more readily apologize and undo mistakes. He theorizes Claude's training makes it assume the human is wrong and needs correcting. The post is personal commentary with no benchmarks or systematic evaluation.