Lip sync looks fake? Check teeth, profile, timing and seams on Sume
sync. labs names four tells of fake lip sync. Pull PNG stills from a Sume avatar clip with video frames at the moments each tell shows up, then decide.
To check an AI avatar clip for fake-looking lip sync, name the tell, find the second it would show, and look at a lossless still from that second. On Sume, POST /v1/video-frames returns up to 24 PNG stills from one media.sume.com clip at times you choose, so you can inspect teeth, head angle and scene joins instead of scrubbing by eye.
The four tells come from the sync. labs blog, which lists a July 14, 2026 post called "Why AI lip sync still looks fake" and summarises the artifacts as "floating teeth, broken side profiles, drifting timing, and a visible seam." The full post was not readable when this page was written, so this guide uses only that list. The Sume side comes from the video frames guide and the avatar video guide.
Which Sume tool finds which tell?
Each tell needs a different kind of evidence, and only some of it is a still.
| Tell | What to pull | Sume route |
|---|---|---|
| Floating teeth | PNG stills on open-mouth syllables | POST /v1/video-frames with format png |
| Broken side profile | Stills where the head turns | POST /v1/video-frames at the turn times |
| Drifting timing | Word times against mouth shapes | video inspect with transcribe true |
| Visible seam | Stills either side of a scene join | POST /v1/video-frames around each boundary |
How do you pull the stills?
Pass exactly one of at[] (1 to 24 values) or fps, and ask for png so compression does not hide an edge. Every time must be at least 0 and below the clip duration, or the job fails with frame_time_out_of_range and names the probed duration. The clip has to be a media.sume.com artifact in your workspace, which an avatar video result already is.
The sketch below reads three stills around a scene join at about 3 seconds and one at 9 seconds, then fetches the resource. Submit is always a 202, so poll the GET until resource_status is ready.
import json
import os
import time
import requests
BASE = "https://api.sume.com/v1/video-frames"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
video_url = os.environ["AVATAR_VIDEO_URL"] # media.sume.com result URL
body = {"video_url": video_url, "at": [2.8, 3.0, 3.2, 9.0], "format": "png"}
sub = requests.post(
BASE,
headers={**H, "Idempotency-Key": "seam-check-001"},
json=body,
timeout=30,
)
sub.raise_for_status()
payload = sub.json()
rid = payload.get("request_id") or payload["data"]["request_id"]
while True:
got = requests.get(f"{BASE}/{rid}", headers=H, timeout=30).json()
text = json.dumps(got)
if '"ready"' in text or '"failed"' in text:
print(text[:500])
break
time.sleep(3)What about timing drift?
Stills cannot show timing. For drift, video inspect can probe the clip, sample stills and, with transcribe: true, return a transcript. Compare word times with the stills at those moments. The longer walkthrough is in check avatar lip sync by transcript word times.
Both routes are billed by their own Modal compute, so take few stills at the moments that matter rather than sampling everything.
What can you do if a clip fails the check?
You cannot patch frames. Change an input and render again: a different scene direction, a cleaner reference photo for the avatar, or a shorter script. A multi-scene video with a shared scene is the place seams are most likely to appear, so a single script clip is a fair control.
Sume does not promise a clip will pass these checks, and a still is not a verdict. It is a way to look at the exact frame instead of arguing about it.
How many stills are enough?
Start with a handful at known moments rather than a dense sweep. at[] takes 1 to 24 values and fps is capped at 2 and 24 frames per call, so one call covers a short clip. Source clips are limited to 300 seconds, which is far above the 60-second ceiling of an avatar video.
A failed instant comes back with a null url and does not fail the whole job, so check each frame in the response rather than assuming all were produced.
Sources
Related posts
More in Sume Avatar 1.0
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
- Shortest AI avatar video: 4 seconds, from $0.74 on Sume Standard
Sume avatar videos run 4 to 60 seconds. Per-second rates for Standard, Plus and Max, the 4 second floor, and what 15, 30 and 60 second clips cost.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume