LiveKit avatar join latency and playback latency: what to measure

LiveKit avatar plugins emit join latency and playback latency. Learn what each means before you pick a live avatar, and what a rendered clip has instead.

5 min readSume
All posts

LiveKit avatar plugins automatically emit two numbers: join latency, the time for the avatar to join the room and publish video, and playback latency, the delay between your agent sending audio and the avatar starting playback. Measure both on your own network before choosing a provider. A rendered Sume clip has neither, because nothing streams. Its equivalent is the time from submit to a completed job.

What LiveKit documents

LiveKit's overview says an avatar worker joins the room as a second participant. Your agent sends audio to the worker instead of publishing it directly, and the worker publishes synchronized audio and video tracks back to users. In a frontend you tell the two apart by participant kind agent and the lk.publish_on_behalf attribute, which is null for the main agent and holds the agent identity for the worker.

The page lists 16 avatar providers, including Anam, D-ID, Runway, Synthesia and Tavus. Only some of them offer Node.js support, so check your language before you commit.

What the two numbers tell you

Join latency is a start-up cost. It shows up once per call, when the face first appears, and decides how long a caller stares at an empty tile.

Playback latency recurs on every turn. It adds to the time your speech model already needs, so a slow avatar makes a fast model feel slow. Track the median and the worst case, since the worst case is what a viewer remembers.

The rendered-clip equivalent

A Sume job moves through queued, processing, then completed, failed or canceled. queued is a normal accepted state, because workspace concurrency limits apply when workers move jobs into processing. Poll GET /v1/jobs/:id/status with exponential backoff, and never resubmit a paid request just because your local process timed out.

So the numbers to log are submit-to-queued-exit and submit-to-completed. They are throughput numbers, not interaction numbers.

Latency vocabulary, live vs rendered (read 2026-10-05)
ConcernLiveKit avatar pluginSume job
Start-upJoin latencyTime in queued until processing
Per turnPlayback latencyNot applicable, the clip is already a file
Done signalTracks published in the roomcompleted status, then /result
FailureParticipant never joinsfailed with a public error; reservation refunded
Retry ruleReconnect the sessionPoll with backoff; reuse the Idempotency-Key

Decision rule

If a human types or speaks and waits for the face to answer, optimise playback latency and choose a live provider. If you control the script, accept job latency and ship a file. For the wider comparison see LiveKit's 16 providers versus a rendered Sume avatar.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume