Pika Speech: 5-minute requests, 5-second clones vs Sume TTS

Pika Speech lists 5 minutes per request, 48 kHz output and a 5-second clone at $0.01 a minute. What that means next to Sume's TTS Router and its limits.

5 min readSume
All posts

Pika Speech is a text-to-speech model that Pika lists at $0.01 per minute, with up to 5 minutes of speech per request, 48 kHz output and voice cloning from five seconds of reference audio. Sume does not list Pika Speech; its text-to-speech route runs Cartesia Sonic models, bills per character, and takes a voice from a Sume avatar rather than an uploaded clip.

This page compares the two on the numbers each side publishes, so you can decide whether a clone-from-five-seconds workflow fits the route you already use. Pika's figures come from its own announcement; Sume's come from the API contract and docs.

What does Pika say Pika Speech does?

Pika describes Pika Speech as a 3-billion-parameter flow-matching diffusion transformer that generates 48 kHz speech in a few denoising steps. The page gives a real-time factor of 0.02 for 3-minute requests, which Pika frames as roughly 1.2 seconds of generation per minute of audio.

Cloning needs five seconds of reference audio, and the page says style, pace and duration are controllable. Training data is described as English and Chinese, 403,000 hours in total. Access is through the Pika API Club.

Two caveats about reading those numbers. The 9x cost claim against ElevenLabs v3 is Pika's own comparison, and the quality metrics (word error rate, speaker similarity) are Pika's own benchmark runs. Neither is an independent test, so treat them as vendor claims and run your own script through it.

How do the limits line up?

Both services return audio asynchronously for long scripts, but they measure the cap differently: Pika in minutes of speech per request, Sume in characters per request plus a ceiling on synthesized audio length.

Pika Speech vs Sume TTS, limits as published (read 2026-10-02)
ItemPika SpeechSume TTS Router
Per-request capUp to 5 minutes of speech20,000 characters; audio over 1,200 s fails with tts_duration_exceeded
Billing unit$0.01 per minute (Pika's listed rate)Per transcript character, provider list x 1.25
Output48 kHz speechmp3 44.1 kHz / 128 kbps by default; wav pcm_s16le on request
Voice cloningFive seconds of reference audioNo clone upload on the request; voice comes from a ready Sume avatar or voice id
ModelsOne model, Pika SpeechCartesia Sonic ids from GET /v1/tts-router/models

Can Sume clone a voice from five seconds of audio?

Not on the TTS request. The Sume TTS contract asks for a transcript plus a voice selector: an avatar id or handle, or a voice id. The documented way to find a usable voice is to list avatars and pick one whose voice status is ready. There is no field for uploading a reference clip on that request.

That is a real gap if your workflow is clone-then-narrate with a fresh sample every time. If your voices are set up once per presenter, the avatar-based selector is enough, and the same voice then drives avatar video. Whichever service you use, keep your own record of consent for any cloned voice.

When does the per-minute vs per-character difference matter?

Per-minute pricing is easy to budget for narration; per-character pricing is easy to budget when you already have the script. A 10-minute read at Pika's listed rate is $0.10 by simple multiplication. For Sume, count the characters in the script (spaces and punctuation count), divide by 1,000 and multiply by the rate shown on the catalog.

Check the live rate with the catalog before you commit, because Sume's price is derived from the provider list at a fixed 1.25 multiplier and the model catalog lists per-model capabilities:

curl https://api.sume.com/v1/tts-router/models \
  -H "Authorization: Bearer $SUME_API_KEY"

What should you test before switching?

Run the same 90-second script through both and listen for names, numbers and Korean or other non-English lines, since Pika lists English and Chinese training data. Then check how each handles a script above its cap.

  • On Pika, split anything over 5 minutes into separate requests.
  • On Sume, keep each request under 20,000 characters and under 1,200 seconds of audio, and use wav when the file will be joined again.
  • Join parts with the Sume timeline audio concat job, which is a flat per-job price with no re-synthesis: see Timeline audio.
  • Confirm usage terms for synthetic voices before you publish.

Sources

Related posts

More in Models

All Models posts

Written by Sume