SpaceXAI just launched Grok Voice Think Fast 2.0 - its most capable speech-to-speech model yet
The numbers are so impressive:
• 82.9% overall Speech-to-Speech Quality Index
• 97.2% on speech reasoning
• 56.5% on agentic voice tasks
• Just 0.70 seconds to first audio
• Uses only 0.4× the reasoning tokens of Think Fast 1.0
Grok Voice Think Fast 2.0 now ranks ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash on overall speech-to-speech benchmarks
But the biggest improvement is real-world listening
Across thousands of short phrases in 24 languages, it delivers 1.5–2× better transcription accuracy than Deepgram Nova 3 and ElevenLabs Scribe v2
In noisy environments and compressed phone calls, that advantage grows to around 10×
It can also reason while speaking, allowing tool calls to begin before it even finishes its first sentence.....with no added latency
Conversations now feel much more natural:
• Shorter responses
• One question at a time
• Less filler
• More fluid dialogue
• Better guidance through complex tasks
SpaceXAI has already deployed it on Starlink's phone line and saw significant increases in both sales conversion and customer support containment
Grok Voice Think Fast 2.0 becomes the default grok-voice-latest model on August 5
Pricing: $0.08 per minute of audio
Voice AI is rapidly evolving from scripted assistants into intelligent agents that can listen, reason, speak, and take action in real time
229 likes 9K views
Linus ✦ Ekenstam
@LinusEkenstam
You're not ready for this 🚨
AI voices can now laugh, whisper and sigh on command
Fish Audio just dropped S2.1 Pro, the most expressive voice AI I've heard. More emotional range than ElevenLabs and Cartesia
6x more affordable than ElevenLabs 💸
Here's the full breakdown 👇
For years, voice AI had one goal: sound human. That race is over. Everything sounds realistic now.
The new benchmark is expressiveness. Emotion, nuance, timing. Can a voice *perform* a line, not just read it?
That's what S2.1 Pro was built for.
I ran the same emotional script through Fish Audio, ElevenLabs and Cartesia. Same words, completely different performances. The pauses, the intonation, the way Fish Audio's voice breathes between lines. It's not close. Watch the video below.
The wild part is how you control it. You direct it like a voice actor, in plain text:
[nervous laugh] → expressive laughing.
[cry] → it cries
[long sigh] → it expels a long sigh
Plus word-level control over pronunciation, emphasis and pacing. Generating speech is out. Directing performances is in.
And it's not just for voiceovers. S2.1 Pro streams at ~90ms to first audio. That's fast enough for live conversation, with rhythm that survives interruptions and topic changes. It's what makes voice agents finally feel human.
The practical stuff:
• Clone a voice from 15 seconds of audio
• 80+ languages in one model
• Open-source roots, open weights models, self-hostable
• 1/6 of the cost of ElevenLabs
HeyGen already integrated Fish Audio. More will follow.
Try it: https://t.co/DugUDKVjfm (there's a free tier)
Reply with a line you'd love to hear an AI voice truly perform, and I'll run the best ones through S2.1 Pro and post the results 👇
347 likes 304.2K views
THE FOUNDERS OF @ELEVENLABS HAVE TO BE SWEATING RIGHT NOW
For years, they absolutely dominated TTS because no open-source options could actually sound human.
That era is officially over.
@FishAudio just dropped their S2.1 Pro model, stepping up as one of the rare voice AI companies offering open-weight models.
Devs now get now natural-language control over emotion, pacing, and delivery, alongside serious real-time performance:
The specs speak for themselves:
→ 2x faster than Cartesia
→ 1/6th the cost of ElevenLabs
→ 56.3 chars/s throughput, ahead of GPT-Realtime-2
→ Native support for 83+ languages
→ Around 90ms latency for natural conversations
Need to keep everything in-house? S2.1 Pro also supports on-prem deployment with zero data retention.
If you're building production-grade voice agents, this is THE most interesting open-source releases to test right now 👀
61 likes 7.2K views