Speech-to-speech vs cascade: does 42% lower latency matter for video?

A newsletter reports speech-to-speech models run 42% lower latency than cascades. Why that helps live agents, not a narrated video, and what to measure.

4 min readSume
All posts

It does not matter for a narrated video, and it matters a lot for a live agent. The Livetok weekly update of October 5 reports that GPT-Live-1 speech-to-speech models show 42% lower latency than cascaded stacks (speech-to-text, then a language model, then text-to-speech). A video voiceover is made before anyone watches, so waiting a few more seconds costs nothing.

What does matter for narration is consistency, pronunciation control and the cost per finished minute. Those are properties of the TTS step, which a cascade lets you pin and test on its own.

What the reports say

The numbers below are as reported in newsletters, not measured by Sume.

  • Neither newsletter links a method, so treat the numbers as a direction, not a benchmark.
Speech-to-speech reports this week (read 2026-10-08)
ItemReported figureSource
GPT-Live-1 speech-to-speech vs cascade42% lower latencyLivetok, Oct 5
Deepslate speech-to-speech440 msKrisp Voice AI newsletter
Mini realtime models (OpenAI)$10 / $20 per 1M audio tokens in / outOpenAI pricing page

Decide in four steps

  • If a person speaks to the system in real time, shortlist speech-to-speech models.
  • If you write a script first, use a cascade: your text, then TTS.
  • Pin the TTS model and voice so the output is repeatable.
  • Check each take with a transcript round trip before you publish.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-router/generate",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "tts-demo-001",
    },
    json={
        "model": "sonic-3.6",
        "transcript": "Welcome back. Today we compare three prices.",
        "voice": {"id": os.environ["SUME_VOICE_ID"]},
        "timestamps": {"words": True},
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

Why the cascade is better for narration

In a cascade, you can edit the text between steps. You can fix a name, change a number, or cut a line, and regenerate that one line as its own job. A speech-to-speech model gives you less to edit, because the words are produced inside the model. With Sume, a TTS job takes the exact transcript you send, up to 20,000 characters, so the text you approved is the text that is spoken.

What to measure instead

For a narrated video, three numbers decide quality: the word error rate of a round trip (speak, then transcribe and diff against the script), the loudness match across takes, and the cost per finished minute. Latency to first audio is not on the list.

Run the round trip on every take. A voice engine can drop a word or misread a figure while still sounding natural, and a viewer hears an error as a trust problem. Our transcription diff recipe shows the check at about four cents a take on Sume.

What Sume does not do

Sume does not offer speech-to-speech models, live conversation or streaming TTS. Sume TTS covers Cartesia Sonic only. The router has no Eleven, OpenAI, Gemini, MAI or Inworld engines, no streaming TTS, and no routing presets. Jobs are asynchronous, and any audio over 1,200 seconds fails with tts_duration_exceeded. I could not verify the 42% or the 440 ms figures beyond the newsletter text.

Related

For a TTS check, see transcribing a take to diff the script and OpenAI realtime pricing against narration.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume