LiveTranslate 2.3 s latency vs STT plus TTS
Qwen3.8-LiveTranslate cut average lagging from 2.8 to 2.3 seconds. Sume has no live interpreter, but stt_create and tts_create can build an offline dub.

Qwen3.8-LiveTranslate, listed on September 18, 2026, uses an interleave architecture and reports average lagging (LAAL) down from 2.8 to 2.3 seconds. Sume does not offer live interpretation; for recorded video you can chain stt_create, a translation step you choose, and tts_create.
What the release note says
The Releasebot Qwen feed records two facts for this model: the interleave architecture and the drop in average lagging (LAAL) from 2.8 s to 2.3 s. A 0.5 second reduction is about 18 percent of the earlier figure. The entry does not say how latency was measured, so compare only against your own tests.
| Item | Value |
|---|---|
| Date listed | Sep 18, 2026 |
| Architecture | Interleave |
| Average lagging (LAAL) before | 2.8 s |
| Average lagging (LAAL) now | 2.3 s |
Live versus offline
Live translation is a streaming problem: audio goes in while the speaker is still talking and the delay is what the listener feels. A dubbing job is a batch problem: you have the whole file, so quality and timing matter more than the first-word delay.
Sume's hosted tool list contains speech-to-text (stt_create) and text-to-speech (tts_create), both submitted as paid jobs with an idempotency_key. There is no live interpreter among the documented tools, so a latency comparison with LiveTranslate would not be meaningful.
An offline dub pipeline
Each step is a separate job, so you can inspect the text between steps.
- Transcribe the source audio with
stt_createand keep the timed text. - Translate the text with the model or service you trust, and have a speaker of the target language review it.
- Generate the voice track with
tts_create, one job per sentence or paragraph. - Use
jobs_waitto collect results, then place the audio on the video with the timeline tools or Sume's other media tools. - Check duration drift: translated speech is often longer or shorter than the source.
When to stay with a live model
If the use case is a meeting, a call or a stream, choose a live model and evaluate its latency in your network. If the use case is a published video, the offline route gives you a reviewable transcript and a file you can re-render.
Sources
Related posts
More in Comparisons
- Capacity fallbacks: Replicate's model swap vs pinning on Sume
Replicate lists a model falling back to another at capacity. On Sume you pin a model or send sume/auto, and a capacity error is retried with the same key.
- Replicate FLUX 1.1 pro, Recraft V3 and Ideogram V3 prices vs Sume rows
Replicate lists FLUX 1.1 pro and Recraft V3 at $0.04 and Ideogram V3 quality at $0.09. Sume rows: $0.0375, $0.05, $0.075. 100 images each, compared.
- Prediction deadlines vs a Sume client deadline: stopping isn't cancel
A client-side deadline stops you watching a Sume job; it does not cancel it. Cancel works only before generation starts, so a missed deadline still bills.
- Runway enterprise exception requests vs Sume's queue-first admission
Runway's docs mention enterprise exception requests for higher volume. Sume takes extra jobs as queued and sets concurrency by plan. How to plan a big batch.
Written by Sume