Grok Voice Transcribe 2.0
xAI releases Grok Voice Transcribe 2.0, claiming top streaming transcription accuracy, twice v1.0's accuracy at unchanged pricing.
xAI launched Grok Voice Transcribe 2.0, a speech-to-text model built on the audio foundation model behind Grok Voice, claiming it ranks first in accuracy among 32 streaming models on the Artificial Analysis leaderboard. Internal evaluations show short-phrase word error rate dropping from 20.6% to 6.8%, with multilingual transcription the largest improvement over version 1.0. Features include batch and streaming modes, word-level timestamps, speaker diarization, up to 8-channel transcription, and key-term biasing. Pricing is unchanged at $0.10 per hour batch and $0.20 per hour streaming; version 2.0 becomes the Speech-to-Text API default while 1.0 will be deprecated, and Atlassian is adopting it for Loom transcription.