Avatar Face Swap (Beta) for ad variants: 4 to 15 second source clips
Sume's Avatar Face Swap Beta puts a ready avatar face on a public source video. What it requires, what it does not take, and how to poll the job.
To put a ready avatar's face onto an existing clip, call POST /v1/models/sume/avatar-face-swap/v1.0/runs with avatar_handle, video_url and quality. The endpoint is a Beta. The source must be a public HTTPS video, and the current plan is for clips of about 4 to 15 seconds with usable audio. It is the right tool when you already have a performance you like and want the same ad with a different face. It is the wrong tool when you want an avatar to say new words, which is the job of the talking-video route.
What the request takes
The required fields are avatar_handle, video_url and quality. In the Beta, quality has no default: send standard, plus or max, or the request is not valid. By design, the endpoint does not support prompts, transcripts, duration knobs, aspect ratio, avatar ids in the body or provider fields. If you need to change the script, make a new talking video instead.
curl -X POST https://api.sume.com/v1/models/sume/avatar-face-swap/v1.0/runs \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: faceswap-ad-017" \
-d '{
"avatar_handle": "your_avatar_handle",
"video_url": "https://example.com/source-ad.mp4",
"quality": "plus"
}'Source video rules
The source must be fetchable. The endpoint rejects localhost, private-network, non-HTTPS, signed or private URLs and provider task URLs. So a pre-signed storage link that carries a signature in the query string is a likely failure, and a plain public HTTPS address is the safe choice. The worker checks that the clip suits the swap, and the docs give 4 to 15 seconds with audio as the current plan, not as a published guarantee.
| Field | Required | Notes |
|---|---|---|
| avatar_handle | Yes | A ready Avatar 1.0 identity |
| video_url | Yes | Public HTTPS video; no signed or private URLs |
| quality | Yes in Beta | standard, plus or max; no default |
| prompt, transcript, duration, aspect ratio | Not supported | Rejected by design |
Poll and read the result
The call is job-backed. Poll GET /v1/jobs/{job_id}/status, then read GET /v1/jobs/{job_id}/result. A finished resource exposes a public-safe video_url under media.sume.com. Use resource_status as the main signal for readiness and job_status when you poll. The endpoint accepts the same communication modes as other generation submits: async by default, sync or subscribe with wait_timeout_seconds, and webhook with a public HTTPS webhook_url.
For an ad batch, send one request per source clip, each with its own idempotency key built from the clip id and the avatar. A retry with the same key and body does not start a second paid job.
Use it with permission
A face swap changes who appears in an ad. Use avatars that you created from your own or a licensed person's photos and consent, and do not apply an avatar to a clip of someone who has not agreed. Label the result as your platform requires. The docs describe the technical limits only, so rights and disclosure are on you.
A practical workflow for variants: keep one source clip per concept, and make one swap per avatar. The cost per ad is then one job, with its own quality tier, so test standard on the first clip, look at the mouth and the jawline against the source audio, and move to plus or max only for the ones you will run. Because the endpoint takes no prompt, the result is decided by the source clip, the avatar and the tier, so improve those three, not the request body.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar inline captions vs standalone video captions: what is billed
Inline captions on a Sume avatar video create no separate caption job. Standalone captions are a billed job on a public URL. When each one is right.
- Avatar video aspect ratios: five choices, 720p only, 9:16 default
Sume Avatar 1.0 accepts 1:1, 3:4, 9:16, 4:3 and 16:9, defaults to 9:16, and renders 720p only. What each choice means for a vertical or landscape ad.
- Avatar package with captions and soundtrack: which video_url you get
In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.
- Avatar video with a product image: the premium is 1-3 cents a second
A product_image on a Sume Avatar 1.0 video adds $0.010 (standard), $0.013 (plus) or $0.030 (max) per second: 30 to 90 cents on a 30-second ad.
Written by Sume