Griffin ranks second on LSE-C: how to judge lip sync yourself
Tavus says LSE-C rewards pronounced mouth movement beyond naturalness. A way to check a Sume avatar clip with timed stills and a word-timed transcript.
What the benchmark footnote says
Tavus's Griffin page, read 2026-10-03, lists Griffin first on DOVER, FID and THEval and second on the lip-sync metric LSE-C, with a score of 7.27. Tavus explains the second place: it says LSE-C rewards pronounced mouth movement beyond what looks natural.
That is a useful caution for any buyer. A metric that scores sync can favour exaggerated mouths, so a high number does not prove that a face looks right. Your own eyes, on your own clip, still decide.
A check that uses timed stills
Sume's video inspect returns a word-timed transcript when transcribe: true and stills at times you name with frames.at, up to 24 per call. The transcript's words[] give each word's time, so you can pull a still at the start of a plosive such as p, b or m and see whether the lips are closed there.
First import the clip with POST /v1/media-imports if it is not already on media.sume.com. Run the transcript once, pick four or five words, then run a second inspect with frames.at set to those times.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lipsync-check-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/avatar.mp4",
"frames": {"at": [0.4, 1.1, 2.3, 3.0], "format": "png"},
"transcribe": true,
"language_code": "en"
}'What to look at
- Lips closed on p, b and m sounds, open on vowels.
- Teeth that stay the same shape from frame to frame.
- The mouth not opening wider than the word needs; an over-wide mouth is what a sync metric can reward.
- Frames near cuts in a multi-scene clip, where a seam can show.
Cost and limits
The transcript is billed at a public $0.01 per audio minute on top of the inspect's compute; the docs tell you to confirm it live in the catalog. A silent clip returns inspect_source_has_no_audio, so check probe.has_audio first. An at value past the clip's length returns frame_time_out_of_range.
A small routine
Run the check on the first clip of a series, not every clip. If the avatar passes on plosives and the mouth shape looks steady, the same avatar and quality tier will behave similarly on later clips. Re-check when you change tier, language or aspect ratio.
Write down what you saw next to the job id. A note such as "closed on b at 1.1 s, teeth steady" is more useful in six months than a feeling.
Sources
Related posts
More in Sume Avatar 1.0
- HeyGen lipsync captions are always on; Sume's are opt-in
HeyGen deprecated enable_caption on translation and lipsync and now always returns SRT and VTT. Sume avatar videos burn captions only when you ask for them.
- Holiday avatar ad roster: 3 presenters x 4 scripts, cost by tier
Three reusable avatars and four holiday scripts make 12 clips. Priced on standard, plus and max with a product image; the avatars cost $2.85.
- LemonSlice API: image plus streaming audio vs Sume lip sync
LemonSlice drives a live avatar from an image and streaming audio. For a recorded line, Sume's lip sync takes a still plus an audio_url and returns a file.
- Lip sync audio under 5 seconds: H3 Max rejects it, use Fabric
Sume's MiniMax H3 Max lip sync only accepts 5 to 14.8 seconds of audio and returns invalid_request outside it. A 4.5-second line goes to Fabric instead.
Written by Sume