ZeroHour
Vendor

Cerebras

0 mentions in 7 days · 1 in 30 days · 3 total · first seen · last

Timeline

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

OpenAI unveiled Jalapeno custom inference chip claiming 1.5-1.9x better perf-per-watt than NVIDIA GB200/GB300, deploying in-house by year-end.

At the 37th Hot Chips conference, OpenAI published first benchmark details for its custom Jalapeno inference chip, claiming 1.5-1.9x more work per watt, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher interactive-workload performance versus NVIDIA GB200/GB300, with the 700W-rated part staying at or below 550W in tests. Deployment into OpenAI's own infrastructure begins by year-end, with Gen 2 deep in development and Gen 3 underway. OpenAI also said GPT-Astra and Codex helped write low-level kernels, reportedly 1.5-1.8x faster than human-expert code for selected attention and MoE blocks. Cerebras CS-5, Groq 3 LPX and Apple M6 were also featured at the conference.

Latent Space · 19d agoAI industry

OpenAI’s GPT-5.6 Sol runs up to 14× faster with Ultrafast mode

OpenAI launched GPT-5.6 Sol Ultrafast mode in limited preview, running up to 14x faster at 750 tokens per second via Cerebras inference.

OpenAI's GPT-5.6 Sol Ultrafast mode is available in limited preview through the OpenAI API, delivering up to 14x faster processing and up to 750 output tokens per second, powered by Cerebras under the companies' ultra-low-latency inference partnership. Preview customers are testing it in production for coding, commerce, financial research, and support applications. OpenAI is also using Ultrafast internally for incident response tasks such as log analysis and trace review, and for research workflows with multiple same-day experiment iterations.

Help Net Security · Aug 14, 2026AI industry

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

OpenAI previews Ultrafast, an API service tier running GPT-5.6 Sol up to 14x faster via Cerebras at up to 750 output tokens per second.

OpenAI announced a preview of Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 14 times the speed of standard inference. The tier is powered by Cerebras hardware and delivers up to 750 output tokens per second. The offering targets latency-sensitive developer workloads on OpenAI's API platform.

OpenAI News · Aug 13, 2026AI tools & infra