LiveKit avatar join latency and playback latency: what to measure
LiveKit avatar plugins emit join latency and playback latency. Learn what each means before you pick a live avatar, and what a rendered clip has instead.
LiveKit avatar plugins automatically emit two numbers: join latency, the time for the avatar to join the room and publish video, and playback latency, the delay between your agent sending audio and the avatar starting playback. Measure both on your own network before choosing a provider. A rendered Sume clip has neither, because nothing streams. Its equivalent is the time from submit to a completed job.
What LiveKit documents
LiveKit's overview says an avatar worker joins the room as a second participant. Your agent sends audio to the worker instead of publishing it directly, and the worker publishes synchronized audio and video tracks back to users. In a frontend you tell the two apart by participant kind agent and the lk.publish_on_behalf attribute, which is null for the main agent and holds the agent identity for the worker.
The page lists 16 avatar providers, including Anam, D-ID, Runway, Synthesia and Tavus. Only some of them offer Node.js support, so check your language before you commit.
What the two numbers tell you
Join latency is a start-up cost. It shows up once per call, when the face first appears, and decides how long a caller stares at an empty tile.
Playback latency recurs on every turn. It adds to the time your speech model already needs, so a slow avatar makes a fast model feel slow. Track the median and the worst case, since the worst case is what a viewer remembers.
The rendered-clip equivalent
A Sume job moves through queued, processing, then completed, failed or canceled. queued is a normal accepted state, because workspace concurrency limits apply when workers move jobs into processing. Poll GET /v1/jobs/:id/status with exponential backoff, and never resubmit a paid request just because your local process timed out.
So the numbers to log are submit-to-queued-exit and submit-to-completed. They are throughput numbers, not interaction numbers.
| Concern | LiveKit avatar plugin | Sume job |
|---|---|---|
| Start-up | Join latency | Time in queued until processing |
| Per turn | Playback latency | Not applicable, the clip is already a file |
| Done signal | Tracks published in the room | completed status, then /result |
| Failure | Participant never joins | failed with a public error; reservation refunded |
| Retry rule | Reconnect the session | Poll with backoff; reuse the Idempotency-Key |
Decision rule
If a human types or speaks and waits for the face to answer, optimise playback latency and choose a live provider. If you control the script, accept job latency and ship a file. For the wider comparison see LiveKit's 16 providers versus a rendered Sume avatar.
Sources
Related posts
More in Developers
- LiveKit avatar providers with Node.js support vs a Sume Node job
LiveKit lists eight avatar providers with Node.js plugins and eight Python-only ones. For a clip rather than a live room, Sume works from Node with fetch.
- Load-test Sume job polling with k6: reads have their own budget
A k6 script that polls one finished Sume job from 5 virtual users, and what 4,800 reads a minute on Free means for any 429 you see.
- Log usage.cost for every Sume video job in Python and total a batch
A finished Sume video job returns usage.cost in USD. Log it with the job id and model, and sum it for a batch. Here is a short Python script that does it.
- Log usage.cost from the Sume images response to CSV in Python
POST /v1/images returns usage.cost as the billed USD amount. A Python logger that writes model, image count and cost per image to a CSV for budget reviews.
Written by Sume