Avatar face swap beta limits: 4 to 15 seconds, quality required
Sume Avatar Face Swap is Beta: a public HTTPS source video of about 4-15 seconds with audio, a required quality tier, and no prompt or aspect options.
Sume Avatar Face Swap 1.0 is a Beta endpoint that applies a ready avatar face onto a public HTTPS source video. The source clip should run roughly 4 to 15 seconds and have usable audio, quality is a required field with no default, and the endpoint does not take prompts, transcripts, duration knobs or aspect ratio.
The request
Per Face swap (Beta), read 2026-10-08, the route is POST /v1/models/sume/avatar-face-swap/v1.0/runs. It is not the old consumer /face-swap route.
| Field | Rule |
|---|---|
| avatar_handle | Required; a ready Avatar 1.0 identity |
| video_url | Required; fetchable public HTTPS video |
| quality | Required: standard, plus or max; no default in Beta |
What it rejects
The endpoint rejects localhost, private-network and non-HTTPS URLs, signed or private URLs, and provider task URLs. By design it does not support prompts, transcripts, duration knobs, aspect ratio, avatar ids in the body or provider fields.
- Source length: approximately 4-15 seconds.
- Audio: the clip needs usable audio.
- Visibility: the URL must be reachable without a signature.
- Fields: only avatar_handle, video_url and quality.
Face swap or talking video?
Choose face swap when a real clip already exists and you want the avatar's face on it. Choose Avatar video when you only have a script and want Sume to generate the whole clip. If your source is longer than 15 seconds, face swap is not the right route; the related post on a 20-second person swap covers that case.
Completion works like other generation jobs. The default mode is async, sync or subscribe waits up to the documented bound, and webhook takes a public HTTPS webhook_url. When a completed resource is ready, use resource_status as the primary readiness signal and job_status when you poll.
Checking a source clip before you submit
Trim the clip to the supported length first, and confirm that it has audible speech. Host it on a public HTTPS URL that does not expire mid-job. Avoid signed links, because the endpoint rejects them by design.
Because quality has no default in Beta, a missing field is an error rather than an assumed tier. Pick standard for a quick check, plus for balanced results, and max when quality matters most, and send it explicitly every time.
As a Beta, the contract may change. Read the live OpenAPI at the Sume API reference before you build a pipeline around it, and keep your own tests so you notice a change early.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar script: 33 words = one 12 s clip, 34 words = two 7 s clips
A Sume Avatar 1.0 sentence of 33 words fits one 12-second clip. Add a 34th word and it splits into two 7-second clips. The numbers, with code.
- A 170-word avatar script in 34-word sentences plans 70 s: refused
Why a 170-word script of five 34-word sentences needs 70 seconds in Avatar 1.0, past the 60-second cap, and how to split it into two 35-second videos.
- Avatar video with a product image: the per-second price gap
On Sume Avatar 1.0 a product_image adds $0.010 to $0.030 per second by tier. For a 30-second spokesperson clip that is $0.30 to $0.90. Arithmetic inside.
- Avatar video succeeded but captions failed: what to do on Sume
On Sume a caption-stage failure is soft: the avatar job can still succeed with a clean video_url and captions.status=failed. How to check it and add captions.
Written by Sume