
Pika Audio ships four generative audio models…
Pika's new Soundtrack, Music, SFX, and Speech models target video pipelines and voice products with aggressive unit economics. The bet: efficient inference beats margin on legacy audio APIs.

Pika's new Soundtrack, Music, SFX, and Speech models target video pipelines and voice products with aggressive unit economics. The bet: efficient inference beats margin on legacy audio APIs.

Audio8-ASR-0.1B targets iPhone Neural Engine deployment with Core ML plus ONNX, seven languages, and roughly 200MB runtime memory. Here is when a 0.1B STT stack beats cloud APIs for voice products.

OpenBMB's VoxCPM2 is a 2B tokenizer-free TTS model with Apache 2.0 weights, 30 languages, and LoRA fine-tuning from 5-10 minutes of audio. Here is how it compares to hosted APIs for voice agent projects.

OpenAI and Anthropic upgraded live voice this year, but the real shift is context: voice that reads your docs, uses your frameworks, and pushes back. Here is the checklist I use for client voice workflows and my own thinking sessions.

ChatGPT Voice plus Projects can turn a spoken update into a Markdown draft, a finished PDF, and a team handoff without touching the keyboard. The trick is scaffolding folders once.