MisoTTS open-weights an 8B voice model claiming 110ms latency

Miso Labs released an emotive 8B TTS with RVQ audio tokens and optional audio context. The 110ms number is hosted TTFB on big GPUs. Here’s what matters if you build voice agents.

SaifullahSaifullah
5 min read
MisoTTS open-weights an 8B voice model claiming 110ms latency

Voice agents die on two things: flat delivery and slow first audio. Miso Labs is attacking both with MisoTTS, an 8B open-weights text-to-speech model that conditions on text and prior audio.

Aoden Teo’s launch pitch is blunt: emotive speech that responds in about 110 milliseconds. I care because voice front desks and WhatsApp follow-ups are a big part of the work I do. Latency and tone are not polish. They are the product.

Weights: MisoLabs/MisoTTS. Code: MisoLabsAI/MisoTTS. Writeup: misolabs.ai/blog/miso-tts-8b.

Architecture in plain English

Most TTS stacks hit a vocabulary wall. Speech is wildly diverse (pitch, emotion, accent, pacing). Making a flat token vocabulary large enough explodes the LM head. Miso’s answer is residual vector quantization (RVQ):

  • 32 codebooks x 2048 codes to a huge addressable audio space without a giant single softmax
  • ~7.7B backbone models the text/audio sequence and predicts the first codebook index
  • ~300M decoder fills in the remaining RVQ depth
  • Audio tokenizer: Mimi
  • Inspired by Sesame-style CSM text-to-dialogue transformers

The second design choice matters for conversation: optional audio context. The model can hear how you sound, then answer in a way that tracks tone instead of reading every reply like a press release.

SpecValue
Parameters~8B total
BackboneLlama-3.2-style ~8B
Audio decoder~300M
Codebooks32 x 2048
Max sequence2048
Languages todayEnglish
License directionModified MIT (check the repo before commercial use)
Soft Paper diagram summarizing RVQ codebook expansion for TTS without exploding parameters

About that 110ms number

Read the fine print before you put it on a slide.

The GitHub README is clear: public 110ms chatter refers to hosted production API time-to-first-byte on H100-class hardware, not the unoptimized local path in the open repo. Local inference on a workstation GPU will be slower until you invest in serving.

Miso compares against vendor-ish baselines in materials (ElevenLabs ~700ms, Sesame ~300ms). Treat those as marketing comparisons. Run your own A/B on your scripts, accents, and telephony path.

Soft Paper callout that 110ms TTFB is a hosted claim to verify on your own GPU

What is actually shippable today

Useful now:

  • Open weights for self-hosting and research
  • Emotive, context-conditioned generations for dialogue-style turns
  • Local deployment when audio cannot leave your VPC

Not there yet (per their own notes):

  • Full-duplex / live turn-taking
  • Lightweight CPU inference
  • Guaranteed 110ms without their serving stack
  • Broad multilingual coverage

That half-duplex limit matters for phone agents. Many production voice bots are still turn-based, so you can prototype. Barge-in and natural overlap need more than a single-turn TTS dump.

How I’d evaluate it for a client voice stack

  1. Clone the repo, load weights on a known CUDA box, measure TTFB and realtime factor on your phrases.
  2. Compare against the hosted voice you already use on the same script set (clarity, warmth, stability, hallucination/skipping).
  3. Test audio-context conditioning: whisper vs raised voice vs rushed caller.
  4. Check license terms for commercial telephony before you wire Twilio / WhatsApp.
  5. Plan orchestration separately: VAD, endpointing, LLM, tools, TTS. MisoTTS is the mouth, not the whole receptionist.

Budget GPU memory honestly. An 8B backbone plus decoder is not a Raspberry Pi project. Plan for a real CUDA box or a small cloud GPU if you want interactive demos. CPU paths exist for curiosity. They are not how you hit conversational latency.

from generator import load_miso_8b gen = load_miso_8b(device="cuda", model_path_or_repo_id="MisoLabs/MisoTTS") # Then generate with text + optional audio context per their README

Where it fits next to the usual voice stack

A production voice agent is a pipeline, not a model card:

StageTypical toolsWhat MisoTTS replaces
Telephony / WhatsAppTwilio, Meta Cloud APINothing (transport stays)
VAD / turn takingSilero, vendor streamingNot yet (half-duplex)
STTWhisper-class / vendorNothing
LLM + toolsHosted or localNothing
TTSElevenLabs, Cartesia, OpenAI, etc.Candidate mouth

The win is self-hosting expressive speech when:

  • Compliance wants audio on your GPU
  • You need tone matching from prior caller audio
  • Per-character SaaS TTS cost gets silly at volume

The loss is operational: you now own CUDA drivers, batching, and on-call when the voice box OOMs at 9am.

A note on “most emotive” claims

Every TTS launch claims emotion. I listen for three boring tells:

  1. Does excitement sound like shouting compression, or real pacing?
  2. Do whispered inputs produce whispered-aware replies, or the same bright default?
  3. Do long sentences keep breath and stress, or flatten after 8 seconds?

Miso’s blog samples (sports commentary, casual chat) are aiming at those tells. Your domain audio (clinic intake, logistics dispatch, Urdu/English code-switch if you need it later) will decide whether it ships.

Why this sits in Voice & ops for me

Founders do not ask for “an 8B RVQ transformer.” They ask why callers hang up, why the bot sounds dead, and why the first word takes forever. Open emotive TTS with a serious latency target is the right direction, even if the open repo is not yet the 110ms experience.

If you want a voice front desk or lead-chase agent wired to your CRM without sounding like a 2019 IVR, book a free discovery call.

Share this post

Related posts