Real-time avatar latency budget: Simli stack vs a rendered clip
Simli claims under 300 ms for speech-to-video, but a full agent turn adds STT, LLM and TTS. Add up the budget, then see when a rendered Sume clip fits better.
Simli's own page claims under 300 ms for its speech-to-video step, but it also lists the rest of a typical voice-agent turn: speech-to-text at about 100-500 ms, an LLM at about 250-450 ms and text-to-speech at about 250-1200 ms. Adding those up gives roughly 600 ms before Simli's step and up to 2450 ms with it, so the avatar is the smallest part of the wait. Sume does not do real-time video: it renders a finished clip as an asynchronous job, so there is no per-turn latency to budget.
The component figures are from Simli's home page, read 2026-10-04. They are the vendor's own typical numbers, not a measurement of mine.
The turn budget, added up
The page frames the latency of an agent as a chain of four steps. I add the stated ranges; the lower total leaves Simli's step out because the page gives only an upper bound for it.
| Step | Stated range (ms) | Running low (ms) | Running high (ms) |
|---|---|---|---|
| Speech to text | 100-500 | 100 | 500 |
| LLM | 250-450 | 350 | 950 |
| Text to speech | 250-1200 | 600 | 2150 |
| Speech to video (Simli) | under 300 | 600 (no lower bound stated) | 2450 |
What that means for a conversation
A person waiting for an answer notices a total near 2.5 seconds, and the three upstream steps account for most of it. That is why the avatar step shrinking to a few hundred milliseconds does not by itself make a conversation feel instant. The levers are streaming the LLM and TTS, not the face renderer.
Where a rendered clip has no latency problem
A rendered clip has a different cost: wait time per job, paid once, before anyone watches. The viewer never sees it. For a welcome message, a policy update or a product explainer, you submit the job, take the job.completed webhook or poll GET /v1/jobs/{id}/status, and publish the MP4. See Jobs and results for the four communication modes.
Standard is the fastest tier on Sume and Max is the slowest, so a content pipeline that wants quick turnaround should pick standard for drafts and render the approved script at a higher tier.
curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: latency-draft-001" \
-d '{"avatar_handle": "product_host", "quality": "standard", "script": "Welcome back. Here is what changed this week."}'Choosing between live and rendered
- Pick a live avatar when the viewer must interrupt, ask a follow-up or get an answer you could not write ahead of time.
- Pick a rendered clip when you know the words in advance, want to review them before anyone sees them and want one file you can reuse, caption and share.
- Pick neither for a topic where a wrong spoken answer is costly, until a person has reviewed the script.
Sources
Related posts
More in Comparisons
- Speech Arena Elo 1,319: what the score means
Eleven v4 reportedly leads Artificial Analysis' Provider Voice Arena at Elo 1,319. What an arena Elo tells you and how to run your own blind test.
- Spotify AI covers and remixes: licensed tool vs an original score
Spotify and UMG announced a paid add-on for fan covers and remixes. A licensed remix tool is not an original video score. Here is the Sume route for the second.
- Stable Diffusion alternatives in 2026: open weights or hosted API
SD 3.5 is still Stability's newest flagship image model. FLUX.2 [klein], Qwen-Image 2.0 and Z-Image Turbo are the open options. How to pick between them.
- Subtitles for a recorded video: 100 ms partials or final text?
MAI-Transcribe-2-Streaming sells 100 ms partials for live subtitling. For a recorded video, final text and word timings matter more. Choosing with Sume STT.
Written by Sume