Check a long avatar video for face drift: 24 inspect stills
After joining avatar jobs, sample the face across the whole video. Sume Video inspect returns up to 24 stills at times you pick, and the probe is unbilled.
Call Video inspect with a frames.at list of up to 24 times spread across the video, then look at the stills in order. If the face, hair or clothing changes across them, the parts of your joined video do not match. The probe and stills are unbilled, and the source can be up to 1800 seconds.
Why check at all?
A long avatar video built from separate jobs has seams. HeyGen's Avatar V page claims no identity drift in a 10-minute module, which is a claim about HeyGen's single-pass product. Sume's avatar jobs run up to 60 seconds, so a 10-minute video is a join of parts, and a quick visual check is cheap.
How do I pick the times?
Take one still just after each join and one just before it. For ten 60-second parts, the joins are at 60, 120 and so on up to 540. Two stills per join is 18 stills, under the cap of 24.
| Option | Value |
|---|---|
| frames.at | 1 to 24 times in seconds |
| frames.fps | Greater than 0, at most 2 |
| max_edge | 64 to 2160, default 768 |
| seek | precise or fast |
| Source limit | 1800 s |
What does the call look like?
The times below straddle the first two joins. Set seek to precise so the still lands on the time you asked for.
import os, requests
at = [59.5, 60.5, 119.5, 120.5]
r = requests.post(
"https://api.sume.com/v1/video-inspect",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"]},
json={"video_url": "https://media.sume.com/artifacts/artf_demo/module.mp4",
"frames": {"at": at, "seek": "precise", "max_edge": 768}},
timeout=60,
)
r.raise_for_status()
print(r.json())What counts as drift?
Compare face shape, skin tone, hairline, accessories and background between the stills on each side of a join. A small change in lighting is normal; a different face is not. If a part looks off, re-run only that part with the same avatar handle and settings.
Does this prove the video is fine?
No. Stills show what the frames at those instants look like. They do not show motion between them. Watch the video before delivering it.
How do I record the result?
Save the stills next to the video and note the time of each one. If you reject a part, write down the part number and the reason, so the next run changes something specific, such as the avatar handle or the tier. The stills are image files you can review in any viewer and attach to a ticket.
Sources
Related posts
More in Developers
- Check a TTS take with an STT round trip: flag skipped words in Python
Transcribe your own voiceover with Sume STT and diff it against the script to catch skipped or changed words before you ship. Python sketch included.
- Check a video request against /v1/videos/models before you submit
Duration, resolution, size and seed errors cost a round trip. A short Python validator reads the model catalog and refuses a bad request locally first.
- Check a transparent GPT Image 2.5 PNG for real alpha in Python
A transparent GPT Image 2.5 result can still look opaque. Ask for background transparent as PNG, then check the alpha channel in Python: a 20-line script.
- Check avatar lip sync: transcript word times, then stills
Transcribe an avatar video with Sume Video inspect, take word timestamps from words[], and pull stills at those instants with video_frames to inspect the mouth.
Written by Sume