How to check a face swap result: audio, length and stills
Check a Sume face swap before publishing: confirm the audio stream, compare length with the source, and sample stills with video inspect. Free probe.

Run Sume's video inspect on the finished face swap and look at three things: does the file have an audio stream, is its length close to the source's, and does the face stay the same across stills. The probe and up to 24 stills are free; add a transcript only if you want to compare the words against the source.
Face Swap (Beta) is a queued job that follows the source clip's timing and re-attaches its audio, so a clean result should match the source on both. This page shows the request and what to read in the response, based on the Face swap docs and the OpenAPI reference read 2026-10-03.
What does the inspect call return?
| Part of the call | What it gives you | Cost |
|---|---|---|
| Probe | Duration, size, rotation, fps, codecs, audio, HDR | Unbilled |
| frames | Up to 24 stills as media.sume.com images | Unbilled |
| transcribe: true | Speech with word timings | Reserves the STT 1.0 per-minute rate |
| Source cap | 1,800 seconds per video | None |
| Input | A media.sume.com video owned by your workspace | None |
What is the request?
Send the media.sume.com URL from the face swap's result. Ask for stills at the moments you care about: the first frame, a mid-sentence frame and the last.
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: swap-check-001" \
-d '{
"video_url": "https://media.sume.com/YOUR_FACE_SWAP_RESULT.mp4",
"frames": {"at": [0.2, 3, 6, 9], "max_edge": 768},
"transcribe": false
}'What should I compare?
- Audio: the probe should report an audio stream. A swap whose source had usable speech should too.
- Length: Face Swap follows the source and has no duration field, so the result should be close to the source's length.
- Face: open the stills side by side. The avatar's face, hair and outfit should hold from first frame to last.
- Scene and prop: the room, framing and any held product should match the source.
- Words: with
transcribe: true, the transcript of the result should match the source speech, because the audio is carried across.
How do I read the sync and async answers?
Inspect's default mode waits up to 30 seconds and answers 200 with the finished result, or 202 with a queued job. On a 202, poll the returned status URL and fetch the result when it is ready, as in Jobs and results. Do not resubmit; stills and the probe are free, but a transcript reservation is not.
Inspect reads only videos already on media.sume.com, with no open-internet fetch. The face swap result is, so you can inspect it directly. Your source clip usually is not, so compare against your local copy, or import it first if it is a public TikTok or Instagram URL.
What does it cost to check every swap?
Almost nothing. The probe and the stills are not billed, so a check on every swap costs only the time to look at the pictures. The transcript is the one paid part: it reserves the speech-to-text per-minute rate for the clip length. A swap is at most about 15 seconds of source, so the reservation is small, but skip it unless the words matter, such as a script with a legal claim.
Keep the checks in a habit. Put the inspect call in the same script that submits the swap, save the stills next to the job's result, and only then move the clip on to captions or publishing.
Which stills should I request?
- The first frame, to confirm the swap starts on the avatar and not the source person.
- A frame during fast hand movement, because hands are replaced too.
- A frame with the product or prop in view, since the swap keeps it.
- The last frame, to catch a late drift in the face or an early cut.
What if the result looks wrong?
Re-check the source first: the Beta targets a steady shot of one person with usable audio between about 4 and 15 seconds. If the source was fine, submit again with a fresh idempotency key and a higher quality tier, and keep the failed job's events for support. A job that finishes but shows no video_url is a readiness question; poll resource_status until it says ready before judging the file.
Sources
Related posts
More in Sume Avatar 1.0
- Does Sume face swap keep the original voice, outfit and scene?
Sume's Face Swap (Beta) keeps the source clip's audio, camera and timing but replaces the whole person, not just the face. What stays, what changes, cost.
- How to make an AI UGC ad look less staged with Avatar 1.0
Less-staged AI UGC comes from the first frame: phone-style framing, a casual scene prompt, an approved preview, then the final render. The levers Sume exposes.
- Shortest AI avatar video: 4 seconds, from $0.74 on Sume Standard
Sume avatar videos run 4 to 60 seconds. Per-second rates for Standard, Plus and Max, the 4 second floor, and what 15, 30 and 60 second clips cost.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
Written by Sume