Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning
Kyutai released two open-weight 9B speech models that raise spoken GSM8K accuracy from 27.3% to 77.1% using reinforcement learning.
Kyutai released two open-weight 9B speech-to-speech models, Voice of Reason, fine-tuned from GLM-4-Voice-9B with supervised fine-tuning and reinforcement learning and no separate text LLM. On spoken GSM8K, accuracy rises from 27.3% for the base model to 70.3% for the direct checkpoint and 77.1% for the STITCH variant that inserts silent reasoning chunks. SFT on 150,616 Orca-Math problems reached 61.7%, and group-relative RL ran for 1,500 updates on 16 H100 GPUs. Removing temperature correction collapsed accuracy to 12.3%, while spoken TriviaQA fell from 40.6% to 34.0%.