Tavus Griffin 0.43 s latency vs a Sume async avatar job
Tavus reports 0.43 s audio-to-video latency for Griffin. Sume Avatar 1.0 is a different design: an async job that returns a finished clip. Here is the contrast.
Griffin's 0.43 second figure and a Sume avatar job measure different things. Tavus reports an average audio-to-video latency of 0.43 seconds on H100 GPUs for a live, full-duplex model, so a viewer sees a face react almost as they speak. A Sume avatar video is an asynchronous job: you submit a script, the job runs, and you fetch a finished clip. Latency per turn is not the metric there; time to a finished, reviewable file is.
Use the numbers below to decide which problem you actually have.
What Tavus states
On the Griffin page Tavus lists 720p output, 320 ms chunks and a single reference photo as the input. Griffin-Lite itself is a research preview limited to select testers, so these figures describe a system you cannot call yet.
| Fact | Value |
|---|---|
| Average audio-to-video latency | 0.43 s on H100 GPUs |
| Chunk size | 320 ms |
| Resolution | 720p |
| Input | A single reference photo |
| Customer access | Not available |
How a Sume job behaves
Sume generation is asynchronous by default and returns 202 with a job. Sync and subscribe modes wait at most 30 seconds (wait_timeout_seconds); if the job is not terminal by then, you poll instead of resubmitting. See jobs and results for the status, events and result endpoints.
For a webhook, terminal events only are delivered: job.completed, job.failed and job.canceled, signed with HMAC SHA256 over <timestamp>.<raw_body>. That is a good fit for a pipeline that publishes the clip when it lands.
Choosing by use case
A support agent that must look attentive while a customer speaks needs a live model. A product video, a lesson opener or an ad needs a reviewed file. Avatar 1.0 accepts a script of 4 to 60 seconds estimated duration and returns a mirrored media.sume.com URL, per the avatar video docs.
Do not compare a per-turn latency with a render time. They answer different questions, and neither tells you which model looks better.
- Need sub-second reaction: a live model, when one is available to you.
- Need a file you can approve and publish: a rendered avatar video.
- Need to avoid blocking a request: submit async and use a webhook.
Sources
Related posts
More in Comparisons
- Tavus per-minute video price vs Sume avatar per second
Tavus lists video generation at $1 per minute overage on Starter. Sume avatar video is $11.04 to $33 per minute depending on quality. Dated 2026-10-01.
- TikTok product avatars skip shoes, hats, sunglasses and bracelets
TikTok's Symphony help page lists products its avatars cannot show. What to try for those SKUs, and what Sume's avatar product_image does and does not claim.
- Seedance 2.5 in Symphony takes 50 references: what Sume lists
TikTok says Seedance 2.5 in Symphony takes up to 50 image, video and audio references. Sume's docs give per-model reference types, so check the catalog first.
- Together AI dynamic rate limits (429, 503) vs Sume rate_limited
Together AI publishes no fixed tiers: limits track live capacity and your recent traffic. How its 429 and 503 map to Sume's rate_limited and queue_full.
Written by Sume