Griffin 10 ms audio packets vs Sume TTS: async jobs, no streaming
Tavus Griffin emits speech in packets as small as 10 ms. Sume's TTS Router lists streaming TTS as a non-goal, so audio arrives as a finished file for lip sync.

Griffin streams: Tavus says it generates video in 320 ms chunks at 25 fps and speech in packets as small as 10 ms. Sume does the opposite on purpose. Its TTS Router lists streaming TTS as out of scope, so you submit text, poll a job, and hand the finished audio file to a lip-sync route.
What Tavus lists
On the Griffin page, Tavus gives 0.43 seconds as the average audio-to-video latency on H100 GPUs, 720p output, video in 320 ms chunks at 25 fps, and speech at a 48 kHz sample rate. These are figures for a research preview open to select trusted testers, not a service you can call.
The small packets are what let a model start speaking before it has decided the whole sentence, which is part of interrupting and overlapping naturally.
What Sume's audio path is
The TTS Router is an async, job-based surface. You POST /v1/tts-router/generate with a required model from the Cartesia Sonic catalog, one transcript source and one voice selector, then poll GET /v1/jobs/:id/status and read /result. The docs list streaming TTS among the first ship's non-goals.
That shape suits a rendered clip, because a lip-sync model needs the whole audio file first. For H3 Max lip sync the file must be Sume-hosted, 10 MB or less and 5 to 14.8 seconds long.
| Item | Griffin-Lite | Sume TTS then lip sync |
|---|---|---|
| Audio unit | Packets as small as 10 ms | One finished file |
| Video unit | 320 ms chunks at 25 fps | One MP4 per job |
| Average audio-to-video latency | 0.43 s on H100 (Tavus figure) | Not a live metric; track job completion |
| Output size | 720p | 480p, 768p or 1080p on H3 Max lip sync |
| Access | Select trusted testers | Public API |
Designing around the file
Split a long script into sentences, synthesise each, and render one lip-sync clip per sentence group. Keep each audio file inside the 5 to 14.8 second window, and never re-cut the audio to fit, since a clamped reservation would produce a clip shorter than the speech. For a window-aware split see driving lip-sync clips from TTS sentence segments.
Poll with exponential backoff and stop on completed, failed or canceled. If you need sub-second turn-taking, a file pipeline is the wrong tool.
Sources
Related posts
More in Developers
- Grok Imagine Image 2.0 makes 10 images per request; Sume's n cap
xAI says grok-imagine-image-2.0 returns up to 10 images per request at $0.04 each. Sume's n is 1-10 per request, with a lower per-model cap on x-ai/grok-image.
- H3 Max job: poll or webhook? A signed Python receiver
Poll GET /v1/jobs/:id/status for one clip, use a signed webhook for batches. Python receiver that refuses an empty secret and checks the sume-v1 signature.
- Lip sync for singing: fal can turn off guidance, Sume has no switch
fal's H3 Max Lip Sync is transcription-guided by default and can be switched off for singing or processed audio. Sume's documented body has no such field.
- HEAD-check reference image URLs before an image edit call (Python)
OpenAI caps edit images at 50 MB and Ideogram at 25 MB. A Python pre-flight that checks size, type and https before you send references to Sume.
Written by Sume