Is my clip ready for face swap? Preflight with reference ingest
Face swap wants a 4 to 15 second source with usable audio. Read the clip first with reference ingest purpose face_swap and check duration and audio.silent.

To check a clip before a face swap, run reference ingest with purpose: "face_swap" and read the manifest: you want a source of about 4 to 15 seconds, and audio.silent should not be true because the beta wants usable audio (Face swap (Beta), Reference ingest). The purpose value is stored but not interpreted, so it labels the read rather than changing it.
Checklist from the two docs
| Check | Where to look | Pass |
|---|---|---|
| Length | last shots[].end | 4 to 15 seconds |
| Audio present | audio.silent | false (silent means integrated loudness at or below -60 LUFS) |
| Speech | audio speech presence | present if the clip has dialogue |
| Public HTTPS URL | your hosting | no localhost, private, signed or provider task URLs |
| Has video | error source_no_video_stream | not returned |
Run the read
With any purpose other than reference_remix, speech.allow_billed_stt defaults to false, so the read does not buy a transcript. A face swap does not need one.
curl -X POST https://api.sume.com/v1/reference-ingest \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: swap-preflight-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/source.mp4",
"purpose": "face_swap"
}'If it fails
A clip longer than 300 seconds returns source_too_long_for_reference_ingest; use video inspect for longer files. A file without a video stream returns source_no_video_stream. A silent track means face swap has nothing usable to work with, so add audio or pick another source.
Cost of being wrong
A swap reserves the beta maximum at the tier rate, so a preflight is cheap insurance.
| quality | Reserved for 15 s |
|---|---|
| standard | $2.76 |
| plus | $3.675 |
| max | $8.25 |
Sources
Related posts
More in Media tools
- Join voiceover takes into one gapless wav: timeline audio concat
Timeline audio concat joins up to 20 hosted audio parts sample-exact into one reusable wav for $0.01 per job and returns segment offsets.
- LinkedIn Page video max ratio 2.4:1: crop a 32:9 recording
LinkedIn Pages accept video from 1:2.4 to 2.4:1 and up to 4096x2304. A 5120x1440 ultrawide capture fails both. Crop it with Sume video filter to fit.
- MAI-Voice 24 kHz 160 kbps MP3 vs Sume TTS 44.1 kHz 128 kbps: mixing
Microsoft's MAI-Voice example saves 24 kHz 160 kbps mono MP3; Sume TTS defaults to 44.1 kHz 128 kbps MP3 or WAV. What the numbers mean for mixing and file size.
- Music from a still: image_url conditioning, still $0.125
The Music Router accepts an optional public HTTPS image_url to condition the track on a still. Request example, keeping a score consistent, fixed price.
Written by Sume