A 20-second ad person swap: Recast fits, avatar face swap stops at 15
For a 20-second source clip, Sume's H3 Max Recast takes 5 to 30 seconds and 1 to 4 person photos; the Avatar Face Swap beta takes about 4 to 15 seconds.
If your source clip runs 20 seconds, use H3 Max Recast. Sume lists h3-max-recast for 5 to 30 seconds, replacing the people in a source video with the people in 1 to 4 reference photos, at 768p or 1080p. The Avatar Face Swap beta is planned for source videos of about 4 to 15 seconds, so a 20-second clip is outside it (Face swap (Beta)).
Two routes, two limits
| Route | Source length | Identity from | Output |
|---|---|---|---|
| Avatar Face Swap 1.0 (beta) | about 4 to 15 s, needs usable audio | A ready avatar_handle | Avatar face on your source video |
| H3 Max Recast (h3-max-recast) | 5 to 30 s | 1 to 4 reference photos | 768p or 1080p |
How Recast is called
Recast is a Video generation model. Per the Video Router and Video generation pages, you send the source as video_url, the person photos as reference_image_urls, one photo per person, and an optional prompt. The duration is the source length, so trim a 33-second clip to 30 seconds first.
Pick the resolution at 768p or 1080p. Read capabilities from GET /v1/video-router/models for the exact envelope before you submit.
When the face swap is still the better fit
If the face should be a reusable Sume avatar and the clip fits 15 seconds, the face swap beta is the match. It requires avatar_handle, video_url and an explicit quality, and it has no default for quality. Neither route replaces consent for a real person's likeness.
Sources
Related posts
More in Sume Avatar 1.0
- 60-second avatar explainer: one $15.48 job or two 30-second jobs
A 60-second Sume Avatar 1.0 explainer with a product image is $15.48 on plus, $11.64 standard, $34.80 max. One job or two 30-second jobs, and what changes.
- AI clone from 2 minutes of video or one photo: what each needs
Tavus asks for two minutes of 1080p video and written consent. Sume Avatar 1.0 starts from a prompt, traits or one photo URL. The inputs compared, dated.
- Avatar clip under 4 seconds: add a silence beat or a longer line
Sume avatar scripts must plan to 4 seconds or more. For a one-liner, add a silence beat in video_inputs or lengthen the line instead of padding with filler.
- Talking avatar from an approved TTS file: image-to-video audio_url
To animate an avatar from a voiceover you already approved, send its Sume-hosted audio_url and duration_seconds to Avatar 1.0 image-to-video. Limits inside.
Written by Sume