SpaceXAI Unveils Grok Voice Transcribe 2.0: A Speech-to-Text API Boasting Twice the Accuracy of Its Predecessor, Priced at $0.10 per Hour
SpaceXAI released Grok Voice Transcribe 2.0, a speech-to-text model claiming twice the accuracy of v1.0 at unchanged pricing ($0.10/hour batch, $0.20/hour streaming). Built on the Grok Voice audio foundation model, it ranks first among 32 streaming models on Artificial Analysis, excels at telephony audio, and cuts multilingual WER from 20.6% to 6.8%. Features include word-level timestamps, speaker diarization, up to 8 channels, key term biasing, and smart turn detection. Atlassian Loom has already adopted it for video transcription workflows.
SpaceXAI has rolled out Grok Voice Transcribe 2.0, its latest speech-to-text model, promising double the accuracy of version 1.0 while keeping pricing unchanged. The model is designed for challenging audio conditions—noisy phone lines, overlapping speakers, regional accents, and spoken credentials—and is available now in both batch and real-time streaming modes through the Speech to Text API. For developers building voice agents and transcription pipelines, it arrives as a serious competitor in a crowded market.
What Is Grok Voice Transcribe 2.0?
The new model is built on the same audio foundation model that powers Grok Voice. According to SpaceXAI, Grok Voice already processes tens of thousands of customer-support calls daily, transcribes millions of hours of video narration, and drives the Grok assistant inside Tesla vehicles.
Training relied on live, noisy, multilingual audio captured in varied real-world environments, followed by post-training refinement. Unlike some competitors, SpaceXAI has not released open weights, so the model is available only as a hosted API under the identifier `grok-voice-transcribe-2.0`.
Benchmarks and Performance Claims
SpaceXAI reports a first-place ranking among 32 streaming models on the public Artificial Analysis leaderboard. That benchmark, AA-WER Streaming, uses roughly 8 hours of audio weighted as follows: AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%.
The company also measured word error rate (WER) on four internal datasets drawn from production traffic: telephony recordings (8 kHz English support calls), English conversations with Grok, English credentials (phone numbers, emails, addresses), and short voice-assistant phrases across 19 languages. Version 2.0 outperformed 1.0 on all four, and SpaceXAI claims it beat every model tested on telephony audio. These figures are vendor-reported and have not been independently verified.
Multilingual accuracy is described as the biggest leap over the previous version. On short utterances—which give models little context for language identification—WER dropped from 20.6% to 6.8%, a reduction of roughly 67%. The model auto-detects languages, handles mid-recording language switches in a single pass, and documents written-form formatting for numbers, currencies, and units across 25 languages.
Developer Features and Pricing
All capabilities ship through the same API:
- Batch and streaming: transcribe files and URLs, or stream audio over WebSocket
- Word-level timestamps with confidence scores
- Speaker diarization at no extra cost
- Multichannel support for up to 8 independent channels
- Key term biasing: up to 100 domain terms per request, 50 characters each
- Written-form formatting for numbers, dates, currencies, phone numbers, and emails
- Filler word removal ("um," "uh") enabled by default
- Smart turn detection using an ML model for voice agents
The batch endpoint accepts files up to 500 MB in 12 audio formats; streaming supports Opus at roughly 4 KB/s compared with 48 KB/s for raw 24 kHz PCM.
Pricing matches version 1.0: $0.10 per hour of audio for batch and $0.20 per hour for streaming—about $1.67 and $3.33 per 1,000 minutes respectively, with diarization, timestamps, and key terms included.
Atlassian Loom Adopts the Model
Atlassian Loom has switched to Grok Voice Transcribe 2.0 for transcribing all videos after finding it more accurate than its previous solution. SpaceXAI describes the resulting workflow as record, transcribe, then code: a user records an action plan in Loom, the transcript feeds into Cursor, and Cursor makes the code changes. Sanchan Saxena, SVP of Teamwork Collection at Atlassian, called it "closing the loop from context to code."
What to Watch
Grok Voice Transcribe 2.0 makes a strong claim—2x accuracy at flat pricing—with third-party benchmark support on Artificial Analysis, though internal results remain vendor-reported. Since it is API-only, developers must set `model=grok-voice-transcribe-2.0` explicitly. Whether independent evaluations confirm the accuracy gains, and whether SpaceXAI eventually releases open weights, will be the key developments to follow.
Meta description: SpaceXAI's Grok Voice Transcribe 2.0 claims 2x the accuracy of 1.0 at the same price—$0.10 per audio hour for batch, with streaming and diarization included.
Tags: SpaceXAI, Grok Voice Transcribe 2.0, speech-to-text, transcription API, voice AI
Featured image: Abstract illustration of flowing sound waves converting into structured digital text lines, rendered in cool blue and violet gradients with no people, logos, or text.
Comments
No comments yet. Be the first to comment.