Avatar clip frame rate: Griffin 25 fps vs Fabric vs H3 Max
Tavus says Griffin streams 8-frame latents at 25 fps. Sume's docs say Fabric is 25 fps and H3 Max must be measured. Probe a clip with video-inspect.
Tavus's Griffin page says one latent is 8 frames at 25 fps, which is 320 ms, and Griffin-Lite generates 720p video in 320 ms chunks in real time (read 2026-10-05). Sume's H3 Max lip-sync notes say Fabric is 25 fps and that you must measure H3 Max on the first real clip. POST /v1/video-inspect returns a probe for exactly that.
Why the frame rate matters
If you assemble several talking clips into one timeline, the output frame rate is locked once. Mixing sources at different rates forces a conversion. That is why Sume's packet guidance says to use one lip-sync model for each run: the model sets the assemble output.fps lock.
Griffin's number is a property of a streaming system. You do not receive Griffin files from Sume, so the comparison is about expectation, not interchange.
| Route | Frame rate fact | Source |
|---|---|---|
| Tavus Griffin-Lite | 8 frames per latent at 25 fps, 320 ms chunks | Tavus page |
| Sume Fabric | 25 fps | Sume H3 Max notes |
| Sume H3 Max lip sync | Measure on the first real clip | Sume H3 Max notes |
| Sume talking video (Avatar 1.0) | Probe the result | video-inspect |
Measure your own clip
Import the finished clip into your workspace first, then inspect it. The default mode is sync and waits up to 30 seconds. If it takes longer you get a 202 and poll the job. frames: false skips stills when you only want the probe.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: probe-h3-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false
}'What to do with the result
Read the frame rate from probe, write it next to the model name in your run notes, and keep every clip in that run on the same model. If a later clip goes to the other model because it falls outside the 5 to 14.8 second window, convert deliberately instead of finding the mismatch in the final export.
Sources
Related posts
More in Media tools
- Baby shower video music: a gentle 90-second slideshow bed
Pick a gentle instrumental for a 90-second baby shower photo slideshow: one generation and a two-minute render, $0.325 on Sume.
- Put a 4:5 AI image on a 9:16 canvas with blurred fill in Pillow
Fill the empty bands of a 9:16 canvas with a blurred, darkened copy of the same Sume image, and place the sharp 4:5 original over it. Code and blur settings.
- Burn captions: pick one of script_text, words, cues or segments
Sume video captions accepts one wording source per job: script_text, words, cues or segments. When each fits and which ones skip speech-to-text.
- Camera orbit and zoom transition prompts for Omni first and last frame
Google says Omni 1.1 Flash handles orbits, zoom transitions and loops between two keyframes. Prompt patterns and costs for the Sume API, from $0.19 per draft.
Written by Sume