
SpaceXAI has released grok-voice-transcribe-2.0, a speech-to-text model built for recorded audio and real-time voice applications.
The company says the new model is twice as accurate as Transcribe 1.0 across its real-world evaluations, with its largest reported improvement appearing in short multilingual commands. Separately, the model currently leads Artificial Analysis’ streaming speech-to-text accuracy benchmark.
Pricing remains unchanged at $0.10 per audio hour for batch transcription and $0.20 per hour for streaming. The API includes speaker diarization, word-level timestamps, eight-channel transcription, key-term biasing, text formatting, filler-word removal, and smart turn detection.
Current documentation lists grok-voice-transcribe-2.0 as the API default. Developers who need the previous version must explicitly pin grok-voice-transcribe-1.0.
Why it matters
Reliable transcription is becoming inexpensive infrastructure for voice agents, customer-support systems, meeting tools, media indexing, and voice-driven developer workflows.
The combination of improved noisy-audio performance, streaming support, and unchanged pricing makes Transcribe 2.0 a practical candidate for production testing. Builders should still evaluate it against their own accents, terminology, audio conditions, and privacy requirements—the “twice as accurate” figure comes from SpaceXAI’s internal evaluations, not an independent model-to-model study.
That caution is especially important in sensitive workflows. As we noted in our coverage of AI transcription errors in clinical records, an average benchmark can conceal failures around accents, specialist vocabulary, rapid speech, and uneven audio quality. Voice systems also need clear policies for consent, retention, access, and deletion before recorded conversations become training, evaluation, or operational data.