LiveTranslate 2.3 s latency vs STT plus TTS

Qwen3.8-LiveTranslate cut average lagging from 2.8 to 2.3 seconds. Sume has no live interpreter, but stt_create and tts_create can build an offline dub.

3 min readSume
All posts

Qwen3.8-LiveTranslate, listed on September 18, 2026, uses an interleave architecture and reports average lagging (LAAL) down from 2.8 to 2.3 seconds. Sume does not offer live interpretation; for recorded video you can chain stt_create, a translation step you choose, and tts_create.

What the release note says

The Releasebot Qwen feed records two facts for this model: the interleave architecture and the drop in average lagging (LAAL) from 2.8 s to 2.3 s. A 0.5 second reduction is about 18 percent of the earlier figure. The entry does not say how latency was measured, so compare only against your own tests.

Qwen3.8-LiveTranslate release facts (read 2026-10-03)
ItemValue
Date listedSep 18, 2026
ArchitectureInterleave
Average lagging (LAAL) before2.8 s
Average lagging (LAAL) now2.3 s

Live versus offline

Live translation is a streaming problem: audio goes in while the speaker is still talking and the delay is what the listener feels. A dubbing job is a batch problem: you have the whole file, so quality and timing matter more than the first-word delay.

Sume's hosted tool list contains speech-to-text (stt_create) and text-to-speech (tts_create), both submitted as paid jobs with an idempotency_key. There is no live interpreter among the documented tools, so a latency comparison with LiveTranslate would not be meaningful.

An offline dub pipeline

Each step is a separate job, so you can inspect the text between steps.

  • Transcribe the source audio with stt_create and keep the timed text.
  • Translate the text with the model or service you trust, and have a speaker of the target language review it.
  • Generate the voice track with tts_create, one job per sentence or paragraph.
  • Use jobs_wait to collect results, then place the audio on the video with the timeline tools or Sume's other media tools.
  • Check duration drift: translated speech is often longer or shorter than the source.

When to stay with a live model

If the use case is a meeting, a call or a stream, choose a live model and evaluate its latency in your network. If the use case is a published video, the offline route gives you a reviewable transcript and a file you can re-render.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume