Griffin leads DOVER, FID, THEval: pull frames to judge a clip
Tavus ranks Griffin-Lite first on DOVER, FID and THEval and second on LSE-C. Judge your own avatar clip by pulling exact stills with Sume video frames.
Tavus says Griffin-Lite ranks first on DOVER, FID and THEval and second on LSE-C with 7.27 (Tavus Griffin page, read 2026-10-06). You can't run those metrics on a Sume clip from Sume, because Sume does not ship them. What you can do is pull exact still frames from your rendered clip with Video frames and check the things those metrics stand for by eye: image quality, a face that stays the same, and a mouth that matches the words.
What the page lists
The page ranks Griffin-Lite first on DOVER (a video quality score), FID and THEval, and second on LSE-C, a lip-sync measure. Tavus gives 7.27 for LSE-C. It does not say here how a clip you produce would score, and none of this is a test of a clip you render with Sume.
The manual version with Sume
POST /v1/video-frames takes one Sume-hosted clip and either at[] times or an fps, and returns durable image artifacts at the source size. A submit always returns 202, so poll the job and read the artifacts when it completes.
Pull a still at the first spoken word, one mid-sentence and one on the last word. Look at the teeth and lips for blur, at the hairline for flicker between stills, and at the eyes for a consistent look. This is a spot check, not a benchmark.
- First spoken word: is the mouth open on the right shape?
- Mid-sentence: does the face stay sharp and the same person?
- Last word: does the clip end cleanly, not mid-gesture?
Cost and limits
Video frames bills by its Modal compute rather than a flat price, and reserves a ceiling at submit. Read the page before looping over many clips. Avatar Video itself renders at 720p at this time.
Sources
Related posts
More in Media tools
- HDR phone clip through Sume video trim: exact mode outputs yuv420p
Sume video trim in exact mode re-encodes with libx264 and yuv420p. Keyframe mode is a stream copy. Check the probe before you promise HDR to a client.
- Music bed under narration: 'no vocals' and 'no spoken word' clauses
Sume's Music docs say to end the prompt with 'Instrumental, no vocals' and add 'no spoken word' only when a narrator will talk over the track.
- Kling motion control job done: a signed Python webhook receiver
Submit a Sume Kling 3.0 motion-control job with mode webhook, then verify the HMAC signature in Python, reject an empty secret and answer fast. Runnable code.
- Loop a 30-second music bed under a 3-minute Reel: the body
A short bed can cover a 180-second Reel with soundtrack.loop, duck_db and a fade-out. One Timeline 1.0 body, the 3-minute bill, and the limits to respect.
Written by Sume