Cartesia released Sonic 3.6 in beta, a text-to-speech model that covers 44 languages and tops Artificial Analysis voice leaderboards. The company claims sub-90ms latency and #1 naturalness scores in their Sonic product page .
I pay attention to TTS releases because voice agents live or die on latency and believability. A 200ms gap between user speech end and agent reply is where callers hang up.
What Sonic 3.6 adds
Cartesia positions Sonic as the voice layer for production agents, not audiobook narration. Key capabilities from their launch materials:
| Capability | Why agents care |
|---|---|
| Sub-90ms latency | Keeps duplex conversations feeling live |
| 44 languages | One model for multilingual front desks without chaining vendors |
| Emotion from transcript | Tone follows context without manual SSML hacking |
[laughter] and non-verbal tags | Insert expressions directly in text |
| Voice clone from ~10s audio | Brand consistency for outbound campaigns |
| Localization | Same speaker identity across languages |
Sonic interprets emotional subtext in the transcript by default. You can still steer with tags when you need a laugh or a pause at an exact beat.

Leaderboard context
Artificial Analysis runs public speech arenas comparing vendors on quality and speed. Cartesia has been campaigning hard on both TTS and speech-to-text rankings. Take any vendor leaderboard with salt, but when multiple independent evals agree, it is worth a bake-off on your own audio conditions.
For client work I test with:
- Background HVAC noise (clinics)
- Speakerphone blur (car dealers)
- Code-switching (English + Spanish in one sentence)
Lab scores do not replace those tapes.
Enterprise hooks
Cartesia advertises HIPAA, SOC 2 Type II, GDPR, and PCI postures, plus on-prem/VPC deployment options. That matters when a healthcare receptionist agent cannot send PHI through a random SaaS endpoint.
Competitors like ElevenLabs, OpenAI, and Google all push enterprise tiers. Sonic's bet is state-space model efficiency: quality at latency points others reserve for smaller voices.
Where I'd use it tomorrow
- After-hours clinic triage where warm tone and fast turn-taking reduce hang-ups.
- Outbound sales dialers that need the same cloned voice in English and Spanish without two billing relationships.
- Internal training sims with adversarial prospect personas (Cartesia highlights this use case on their site).
I would not rip out a working ElevenLabs or OpenAI Realtime stack on launch day. I would run Sonic 3.6 on 500 production utterances and measure WER on the return path plus human "sounds robotic" ratings.
Pricing reality check
Cartesia lists plans on their site but pushes larger deployments to sales. Voice margins are sensitive. Sub-90ms often means premium tier or dedicated capacity. Model routing deals (see Stripe/OpenRouter rumors) do not apply here yet. Budget per completed minute, not per character quote alone.
Bottom line
Sonic 3.6 is another proof point that voice agents are a infrastructure category, not a demo trick. Forty-four languages and sub-90ms latency in one model shrinks the excuse list for shipping multilingual voice in production.
Run your own noisy-room eval before you switch. Leaderboards sell meetings. Your callers decide renewals.
If you are scoping a voice agent for ops and want help picking TTS, telephony, and compliance boundaries, book a free discovery call.

