Griffin 10 ms audio packets vs Sume TTS: async jobs, no streaming

Tavus Griffin emits speech in packets as small as 10 ms. Sume's TTS Router lists streaming TTS as a non-goal, so audio arrives as a finished file for lip sync.

5 min readSume
All posts

Griffin streams: Tavus says it generates video in 320 ms chunks at 25 fps and speech in packets as small as 10 ms. Sume does the opposite on purpose. Its TTS Router lists streaming TTS as out of scope, so you submit text, poll a job, and hand the finished audio file to a lip-sync route.

What Tavus lists

On the Griffin page, Tavus gives 0.43 seconds as the average audio-to-video latency on H100 GPUs, 720p output, video in 320 ms chunks at 25 fps, and speech at a 48 kHz sample rate. These are figures for a research preview open to select trusted testers, not a service you can call.

The small packets are what let a model start speaking before it has decided the whole sentence, which is part of interrupting and overlapping naturally.

What Sume's audio path is

The TTS Router is an async, job-based surface. You POST /v1/tts-router/generate with a required model from the Cartesia Sonic catalog, one transcript source and one voice selector, then poll GET /v1/jobs/:id/status and read /result. The docs list streaming TTS among the first ship's non-goals.

That shape suits a rendered clip, because a lip-sync model needs the whole audio file first. For H3 Max lip sync the file must be Sume-hosted, 10 MB or less and 5 to 14.8 seconds long.

Streaming speech vs a finished audio file (read 2026-10-05)
ItemGriffin-LiteSume TTS then lip sync
Audio unitPackets as small as 10 msOne finished file
Video unit320 ms chunks at 25 fpsOne MP4 per job
Average audio-to-video latency0.43 s on H100 (Tavus figure)Not a live metric; track job completion
Output size720p480p, 768p or 1080p on H3 Max lip sync
AccessSelect trusted testersPublic API

Designing around the file

Split a long script into sentences, synthesise each, and render one lip-sync clip per sentence group. Keep each audio file inside the 5 to 14.8 second window, and never re-cut the audio to fit, since a clamped reservation would produce a clip shorter than the speech. For a window-aware split see driving lip-sync clips from TTS sentence segments.

Poll with exponential backoff and stop on completed, failed or canceled. If you need sub-second turn-taking, a file pipeline is the wrong tool.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume