Tavus Griffin 0.43 s latency vs a Sume async avatar job

Tavus reports 0.43 s audio-to-video latency for Griffin. Sume Avatar 1.0 is a different design: an async job that returns a finished clip. Here is the contrast.

4 min readSume
All posts

Griffin's 0.43 second figure and a Sume avatar job measure different things. Tavus reports an average audio-to-video latency of 0.43 seconds on H100 GPUs for a live, full-duplex model, so a viewer sees a face react almost as they speak. A Sume avatar video is an asynchronous job: you submit a script, the job runs, and you fetch a finished clip. Latency per turn is not the metric there; time to a finished, reviewable file is.

Use the numbers below to decide which problem you actually have.

What Tavus states

On the Griffin page Tavus lists 720p output, 320 ms chunks and a single reference photo as the input. Griffin-Lite itself is a research preview limited to select testers, so these figures describe a system you cannot call yet.

Griffin figures from Tavus (read 2026-10-02)
FactValue
Average audio-to-video latency0.43 s on H100 GPUs
Chunk size320 ms
Resolution720p
InputA single reference photo
Customer accessNot available

How a Sume job behaves

Sume generation is asynchronous by default and returns 202 with a job. Sync and subscribe modes wait at most 30 seconds (wait_timeout_seconds); if the job is not terminal by then, you poll instead of resubmitting. See jobs and results for the status, events and result endpoints.

For a webhook, terminal events only are delivered: job.completed, job.failed and job.canceled, signed with HMAC SHA256 over <timestamp>.<raw_body>. That is a good fit for a pipeline that publishes the clip when it lands.

Choosing by use case

A support agent that must look attentive while a customer speaks needs a live model. A product video, a lesson opener or an ad needs a reviewed file. Avatar 1.0 accepts a script of 4 to 60 seconds estimated duration and returns a mirrored media.sume.com URL, per the avatar video docs.

Do not compare a per-turn latency with a render time. They answer different questions, and neither tells you which model looks better.

  • Need sub-second reaction: a live model, when one is available to you.
  • Need a file you can approve and publish: a rendered avatar video.
  • Need to avoid blocking a request: submit async and use a webhook.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume