Pika's new 3B speech model generates a full minute of studio-quality audio in 1.2 seconds, and it could cut your voice pipeline costs by up to nine times
Pika Labs has released Pika Speech, a 3-billion-parameter text-to-speech model running at a real-time factor of 0.02, with 48 kHz output and support for requests up to five minutes long. In internal tests, it outpaces ElevenLabs v3 on cost by a factor of nine.
Pika Labs has been known for its video generation tools, so a text-to-speech model announcement catches you slightly off guard. It should not. For anyone building AI video at any kind of scale, the voice pipeline is often the most annoying and expensive piece, and Pika Speech is a direct challenge to the incumbents who have owned that space.
The headline number is a real-time factor of 0.02, which means the model generates audio roughly fifty times faster than the audio actually plays back. In Pika's own locally-run long-form tests, one minute of 48 kHz studio-quality speech took approximately 1.2 seconds to produce. The model accepts inputs up to five minutes in length, which is meaningfully longer than many competing services cap their requests. On cost, Pika claims up to nine times cheaper than ElevenLabs v3, 4.5 times cheaper than both Cartesia and ElevenLabs Turbo, and twice as cheap as Fish Audio.
What is actually driving the speed
Pika is not being coy about the engineering. Per @pika_labs, the efficiency comes from a full inference stack redesign rather than simply reducing step count. The model uses flow matching combined with distribution-matching distillation, limiting denoising to eight steps or fewer. On top of that: FlashAttention-3 across both self- and cross-attention layers; token packing so that five sentence chunks cost roughly twice as much as one, not five times; and CUDA graph replay with fused RoPE kernels. Each optimisation stacks on the others, and together they brought one-minute generation down from 1.29 seconds to 1.04 seconds in internal benchmarks.
The pace and duration control is also worth understanding properly. Most TTS systems handle timing by trimming or time-stretching audio after generation, which introduces artefacts. Pika Speech places an end-of-speech latent anchor, called an EOS latent, at the target final frame during generation itself. Move that anchor earlier and the delivery compresses naturally. Move it later and the same words relax. This is an inside-model solution, not a post-processing one, and it matters for anything where sync to picture is critical.
What was not said
Pika described these as locally-run long-form tests. That is an important caveat. Latency in production, under API load, across different hardware, may look different. No independent benchmark has been published yet, and claims of nine times cost advantage over ElevenLabs v3 are Pika's own comparison, not a neutral audit. The comparison metrics used and the exact test conditions have not been disclosed in full.
No pricing has been confirmed for production API access. The model is currently accessible through the Pika API Club, which is a separate access tier from Pika's standard video generation subscription. Whether Pika Speech will eventually fold into existing Pika plans or sit as a standalone product is not known. There is also no date given for broader availability beyond the current API Club access.
Language support scope is unspecified in the announcement. The model demonstrates English clearly, but multilingual capability, accent coverage and speaker cloning depth have not been addressed publicly.
Where this fits in an AI filmmaker's pipeline
For working AI video creators, the bottleneck this addresses is real. You have a finished video, you need narration or character dialogue, and the existing options are either expensive at volume or slow enough to break a same-day iteration loop. A model that produces five minutes of speech in well under ten seconds, at competitive audio quality, and billed at a fraction of the current market rate, changes what is feasible in a single session.
The EOS latent approach to duration control is particularly useful for sync-to-picture work. If you are cutting to a beat or a scene transition, having timing baked into generation rather than wrestled in post is a genuine workflow improvement. Test it against your longest and most timing-sensitive scripts first. That is where the difference between inside-model control and post-processing stretch will be most apparent.
What to hold in reserve: wait for an independent latency test under real API conditions before baking Pika Speech into any production pipeline. The numbers are promising and the engineering rationale is coherent, but a single vendor benchmark is not enough to retire your current setup.
---
Sources: announcement thread, efficiency detail thread, EOS latent explanation, technical writeup link, API Club access