Real-time avatar latency budget: Simli stack vs a rendered clip

Simli claims under 300 ms for speech-to-video, but a full agent turn adds STT, LLM and TTS. Add up the budget, then see when a rendered Sume clip fits better.

5 min readSume
All posts

Simli's own page claims under 300 ms for its speech-to-video step, but it also lists the rest of a typical voice-agent turn: speech-to-text at about 100-500 ms, an LLM at about 250-450 ms and text-to-speech at about 250-1200 ms. Adding those up gives roughly 600 ms before Simli's step and up to 2450 ms with it, so the avatar is the smallest part of the wait. Sume does not do real-time video: it renders a finished clip as an asynchronous job, so there is no per-turn latency to budget.

The component figures are from Simli's home page, read 2026-10-04. They are the vendor's own typical numbers, not a measurement of mine.

The turn budget, added up

The page frames the latency of an agent as a chain of four steps. I add the stated ranges; the lower total leaves Simli's step out because the page gives only an upper bound for it.

Simli agent latency components and totals (read 2026-10-04)
StepStated range (ms)Running low (ms)Running high (ms)
Speech to text100-500100500
LLM250-450350950
Text to speech250-12006002150
Speech to video (Simli)under 300600 (no lower bound stated)2450

What that means for a conversation

A person waiting for an answer notices a total near 2.5 seconds, and the three upstream steps account for most of it. That is why the avatar step shrinking to a few hundred milliseconds does not by itself make a conversation feel instant. The levers are streaming the LLM and TTS, not the face renderer.

Where a rendered clip has no latency problem

A rendered clip has a different cost: wait time per job, paid once, before anyone watches. The viewer never sees it. For a welcome message, a policy update or a product explainer, you submit the job, take the job.completed webhook or poll GET /v1/jobs/{id}/status, and publish the MP4. See Jobs and results for the four communication modes.

Standard is the fastest tier on Sume and Max is the slowest, so a content pipeline that wants quick turnaround should pick standard for drafts and render the approved script at a higher tier.

curl -X POST https://api.sume.com/v1/avatar-1.0/talking-video \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: latency-draft-001" \
  -d '{"avatar_handle": "product_host", "quality": "standard", "script": "Welcome back. Here is what changed this week."}'

Choosing between live and rendered

  • Pick a live avatar when the viewer must interrupt, ask a follow-up or get an answer you could not write ahead of time.
  • Pick a rendered clip when you know the words in advance, want to review them before anyone sees them and want one file you can reuse, caption and share.
  • Pick neither for a topic where a wrong spoken answer is costly, until a person has reviewed the script.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume