OpenAI’s GPT-5.6 Sol runs up to 14× faster with Ultrafast mode
OpenAI launched GPT-5.6 Sol Ultrafast mode in limited preview, running up to 14x faster at 750 tokens per second via Cerebras inference.
OpenAI's GPT-5.6 Sol Ultrafast mode is available in limited preview through the OpenAI API, delivering up to 14x faster processing and up to 750 output tokens per second, powered by Cerebras under the companies' ultra-low-latency inference partnership. Preview customers are testing it in production for coding, commerce, financial research, and support applications. OpenAI is also using Ultrafast internally for incident response tasks such as log analysis and trace review, and for research workflows with multiple same-day experiment iterations.