AI face swap video cost: the $2.76 to $8.25 ceiling per clip
Sume's Beta avatar face swap reserves a 15-second maximum: $2.76 on standard, $3.68 on plus, $8.25 on max. Constraints and when to use talking video instead.

A Beta face-swap clip on Sume is estimated at $2.76 on standard, $3.68 on plus and $8.25 on max. Those figures are the contract estimate in the catalog: the Avatar Video no-product rate for the tier times the 15-second Beta source-video maximum. So the ceiling for one swap is priced as if the source were a full 15 seconds, and quality is required; there is no default.
Face swap puts a ready avatar's face onto a short public video. It is not script-driven generation. If you want a presenter to say new words, use the talking-video route.
Request and limits
One endpoint, three required fields.
| Item | Value |
|---|---|
| Endpoint | POST /v1/models/sume/avatar-face-swap/v1.0/runs |
| Required body | avatar_handle, video_url, quality |
| Source video | Public HTTPS, fetchable; about 4 to 15 seconds with usable audio |
| Rejected | Localhost, private network, non-HTTPS, signed or private URLs, provider task URLs |
| Unsupported fields | Prompts, transcripts, duration, aspect ratio, provider fields |
| Status | Beta; prefer resource_status for readiness and job_status for polling |
Standard, plus or max?
The only way to find out what the extra spend buys is to run the same clip at two tiers. Because a swap is capped near 15 seconds, the experiment is cheap: standard and max together are $11.01 at the ceiling. Run both on a hard case, such as a face that turns away from camera or a clip with fast motion, and compare the edges around the jaw and hairline.
curl -X POST https://api.sume.com/v1/models/sume/avatar-face-swap/v1.0/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: swap-demo-standard" \
-d '{
"avatar_handle": "studio_presenter",
"video_url": "https://example.com/inputs/source-video.mp4",
"quality": "standard",
"mode": "async"
}'Use face swap or talking video?
- Face swap: you already have a short performance you want to keep, with its movement and audio.
- Talking video: you have a script and want a presenter to deliver it.
- Neither: you want a long clip. Both stop at short windows (swap about 15 seconds, talking video 60).
Limits and care
This is a Beta surface and its limits can change, so read the OpenAPI before depending on it. Use faces you have the right to use. Platform rules on altered or synthetic people differ, and several platforms ask for a disclosure label on face-swapped content, so check the destination before you publish. Keep the source clip and the avatar handle with the output in your own records, so you can show later what was swapped onto what, and delete test clips you no longer need. A short consent note from the person whose footage you use is cheap to collect and hard to reconstruct afterwards.
Sources
Related posts
More in Sume Avatar 1.0
- avatar-1.0/image-to-video is deprecated: switch to veed/fabric-1.0
Sume's avatar-1.0 image-to-video routes are deprecated aliases of VEED Fabric 1.0. Same body, new URL: audio_url, duration_seconds and one image source.
- Avatar 1.0 talking video: 4 to 60 seconds, cost by quality tier
Work out what a 4, 15, 30 or 60 second Avatar 1.0 talking video costs at standard, plus and max quality, using the per-second rates in the Sume catalog.
- Avatar handle with a leading @: how Sume stores and reuses it
Sume accepts an avatar_handle with or without a leading @ and stores it without. Why a stable handle beats a generated id, and how to reuse it.
- Avatar video soundtrack: send prompt or audio_url, volume 0.05-0.4
Sume's Avatar Video package accepts a soundtrack with exactly one of prompt or audio_url and a volume from 0.05 to 0.4, default 0.15. What each costs and does.
Written by Sume