Check avatar lip sync: transcript word times, then stills

Transcribe an avatar video with Sume Video inspect, take word timestamps from words[], and pull stills at those instants with video_frames to inspect the mouth.

4 min readSume
All posts

Run Video inspect with transcribe set to true, read the words[] timestamps, then call Video frames at three or four of those instants and look at the mouth in each still. This does not measure sync automatically. It gives you stills at known spoken moments so you can judge by eye. HeyGen's Avatar V page claims phoneme-level lip sync across 175-plus languages; Sume's docs make no such claim, so verify the output yourself.

What does transcription give me?

Transcribe is billed at $0.01 per audio minute. The result includes words[] with timings and optional sentence segments. The audio hint is limited to 600 seconds. The probe and any stills in the same call are unbilled.

Steps and limits. Sume Video inspect and Video frames docs, read 2026-10-02.
StepToolCostLimit
Words and timesVideo inspect, transcribe true$0.01 per audio minuteHint max 600 s
Stills at word timesVideo frames, at[]Unbilled1 to 24 values, source up to 300 s

How do I choose the instants?

Pick words with strong mouth shapes, such as a word starting with p, b or m, where the lips close, and a word with an open vowel. Use the start time of each word from words[].

What does the frames call look like?

Video frames is always asynchronous: it returns 202, so poll the job or use jobs wait over MCP. The values below are placeholders for your own word times.

import os, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
     "Idempotency-Key": "lipsync-check-001"}
r = requests.post(
    "https://api.sume.com/v1/video-frames",
    headers=H,
    json={"video_url": "https://media.sume.com/artifacts/artf_demo/avatar.mp4",
          "at": [1.24, 3.8, 6.1], "format": "png", "max_edge": 1024},
    timeout=60,
)
r.raise_for_status()
print(r.status_code, r.json())

What can I conclude?

If the lips are closed on a p or b word and open on a vowel, the sync is plausible at those points. If the mouth is clearly out of step, regenerate that part. Three stills do not cover the whole video, so sample more points for a longer one.

What if the video is over 300 seconds?

Video frames accepts a source of up to 300 seconds. Cut a window with Video trim first, then sample it.

Why use words and not fixed intervals?

Fixed intervals can land on silences, where the mouth is closed and tells you nothing. Word times put each still on speech. They also let you ask for the same sounds in every part, which makes parts easier to compare. If a transcript word is wrong, the time is still usable, because it marks where audio was detected.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume