Speech-to-speech vs cascade: does 42% lower latency matter for video?
A newsletter reports speech-to-speech models run 42% lower latency than cascades. Why that helps live agents, not a narrated video, and what to measure.

It does not matter for a narrated video, and it matters a lot for a live agent. The Livetok weekly update of October 5 reports that GPT-Live-1 speech-to-speech models show 42% lower latency than cascaded stacks (speech-to-text, then a language model, then text-to-speech). A video voiceover is made before anyone watches, so waiting a few more seconds costs nothing.
What does matter for narration is consistency, pronunciation control and the cost per finished minute. Those are properties of the TTS step, which a cascade lets you pin and test on its own.
What the reports say
The numbers below are as reported in newsletters, not measured by Sume.
- Neither newsletter links a method, so treat the numbers as a direction, not a benchmark.
| Item | Reported figure | Source |
|---|---|---|
| GPT-Live-1 speech-to-speech vs cascade | 42% lower latency | Livetok, Oct 5 |
| Deepslate speech-to-speech | 440 ms | Krisp Voice AI newsletter |
| Mini realtime models (OpenAI) | $10 / $20 per 1M audio tokens in / out | OpenAI pricing page |
Decide in four steps
- If a person speaks to the system in real time, shortlist speech-to-speech models.
- If you write a script first, use a cascade: your text, then TTS.
- Pin the TTS model and voice so the output is repeatable.
- Check each take with a transcript round trip before you publish.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-router/generate",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001",
},
json={
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare three prices.",
"voice": {"id": os.environ["SUME_VOICE_ID"]},
"timestamps": {"words": True},
},
timeout=30,
)
r.raise_for_status()
print(r.json())Why the cascade is better for narration
In a cascade, you can edit the text between steps. You can fix a name, change a number, or cut a line, and regenerate that one line as its own job. A speech-to-speech model gives you less to edit, because the words are produced inside the model. With Sume, a TTS job takes the exact transcript you send, up to 20,000 characters, so the text you approved is the text that is spoken.
What to measure instead
For a narrated video, three numbers decide quality: the word error rate of a round trip (speak, then transcribe and diff against the script), the loudness match across takes, and the cost per finished minute. Latency to first audio is not on the list.
Run the round trip on every take. A voice engine can drop a word or misread a figure while still sounding natural, and a viewer hears an error as a trust problem. Our transcription diff recipe shows the check at about four cents a take on Sume.
What Sume does not do
Sume does not offer speech-to-speech models, live conversation or streaming TTS. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I could not verify the 42% or the 440 ms figures beyond the newsletter text.
Related
For a TTS check, see transcribing a take to diff the script and OpenAI realtime pricing against narration.
Sources
Related posts
More in Comparisons
- Speech-to-text per hour: Sume $0.60 batch vs MAI streaming $0.54
Sume STT is batch at $0.01 per audio minute, $0.60 an hour. MAI-Transcribe-2-Streaming lists $0.54 an hour through year end. Different products; the math.
- sume/auto or a pinned image model: what the response shows
With sume/auto, job.model stays sume/auto, the model list omits it and the family is never named. A pinned id echoes back as requested.
- Sume Format vs pasting the same prompt each time: what changes
A Format stores the house style once, then each API run sends only inputs. How the instruction is composed, what a run returns, and the spend cap.
- STT language: Sume language_code hint vs the 60 languages MAI lists
Sume STT takes an optional language_code and auto-detects when omitted. A tracker lists MAI-Transcribe-2-Streaming at 60 languages; Sume's docs state no count.
Written by Sume