Check avatar lip sync: transcript word times, then stills
Transcribe an avatar video with Sume Video inspect, take word timestamps from words[], and pull stills at those instants with video_frames to inspect the mouth.
Run Video inspect with transcribe set to true, read the words[] timestamps, then call Video frames at three or four of those instants and look at the mouth in each still. This does not measure sync automatically. It gives you stills at known spoken moments so you can judge by eye. HeyGen's Avatar V page claims phoneme-level lip sync across 175-plus languages; Sume's docs make no such claim, so verify the output yourself.
What does transcription give me?
Transcribe is billed at $0.01 per audio minute. The result includes words[] with timings and optional sentence segments. The audio hint is limited to 600 seconds. The probe and any stills in the same call are unbilled.
| Step | Tool | Cost | Limit |
|---|---|---|---|
| Words and times | Video inspect, transcribe true | $0.01 per audio minute | Hint max 600 s |
| Stills at word times | Video frames, at[] | Unbilled | 1 to 24 values, source up to 300 s |
How do I choose the instants?
Pick words with strong mouth shapes, such as a word starting with p, b or m, where the lips close, and a word with an open vowel. Use the start time of each word from words[].
What does the frames call look like?
Video frames is always asynchronous: it returns 202, so poll the job or use jobs wait over MCP. The values below are placeholders for your own word times.
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "lipsync-check-001"}
r = requests.post(
"https://api.sume.com/v1/video-frames",
headers=H,
json={"video_url": "https://media.sume.com/artifacts/artf_demo/avatar.mp4",
"at": [1.24, 3.8, 6.1], "format": "png", "max_edge": 1024},
timeout=60,
)
r.raise_for_status()
print(r.status_code, r.json())What can I conclude?
If the lips are closed on a p or b word and open on a vowel, the sync is plausible at those points. If the mouth is clearly out of step, regenerate that part. Three stills do not cover the whole video, so sample more points for a longer one.
What if the video is over 300 seconds?
Video frames accepts a source of up to 300 seconds. Cut a window with Video trim first, then sample it.
Why use words and not fixed intervals?
Fixed intervals can land on silences, where the mouth is closed and tells you nothing. Word times put each still on speech. They also let you ask for the same sounds in every part, which makes parts easier to compare. If a transcript word is wrong, the time is still usable, because it marks where audio was detected.
Sources
Related posts
More in Developers
- Check caption reading speed from STT segments and flag dense sentences
Request sentence segments from Sume STT, compute characters per second for each, and flag the ones too dense to read on screen. Python and caption fixes.
- Check a generated ad image against your brand hex color in Python
The Image API takes a prompt, not a color parameter. Measure each result against your brand hex with Pillow and flag drift before the image goes in an ad.
- China's implicit AI label: provider name and content ID in metadata
China's implicit AI label is metadata with the provider name and a content ID. Why a re-encode can drop it, and how to check a delivered file with ffprobe.
- Choose an image model in code: filter GET /v1/images/models
Pick a Sume image model by what the job needs: references, 4:5, a mask, transparency. A Python filter over the catalog descriptors, with the price lookup.
Written by Sume