Talking video, face swap, Recast or motion control: pick by your input
Four Sume routes make a person on video. Which one fits depends on whether you have a script, a source video, a photo of each person or a motion clip.

Start from what you have
Sume has four routes that put a specific person on screen. They take different inputs and bill differently, so choosing by input saves a wasted call.
| Route | You supply | Limits | Price basis |
|---|---|---|---|
| POST /v1/avatar-1.0/talking-video | avatar_handle and a script | 4 to 60 s estimated | $0.184, $0.245 or $0.55 per second by tier |
| Face Swap (Beta) | avatar_handle, video_url, quality | Source about 4 to 15 s with usable audio | Reserves 15 s at the tier rate |
| Video Router h3-max-recast | 5 to 30 s source and 1 to 4 person photos | 768p or 1080p | $0.375 or $0.5625 per second |
| Kling 3.0 motion control | Image or avatar, plus a motion video | 1 to 30 s | $0.126 per second times 1.25 |
Decision rules
- You have words but no footage: talking-video.
- You have a real video and want one Sume avatar to replace the speaker: Face Swap. It has no prompt, transcript, duration or aspect-ratio field.
- You have a 5 to 30 s clip and photos of each person to put in: Recast. One photo per person, up to four.
- You want a still character to copy the movement from a reference clip: motion control with image_url or an avatar handle, plus motion_video_url.
Worked cost for a 10 second job
Talking video on plus: $2.45. Recast at 768p: 10 s at $0.375 is $3.75; at 1080p, $5.625. Face swap is billed from the reserved source length, so a short source still holds up to 15 s at the tier rate: $3.675 on plus before any release.
Check the model page for your route before you set a budget; rates move.
Common mistakes
Sending a prompt to Face Swap, which has no prompt field. Sending a 40 s source to Recast, which accepts 5 to 30 s. Writing a 90 s script for talking-video, which must estimate to 60 s or less.
Sources
Related posts
More in Models
- Tavus Human Interaction Model: what Griffin is, and what ships today
Tavus calls Griffin a Human Interaction Model. Only select testers have Griffin-Lite. Here is what it is, and what ships today on Sume for a talking presenter.
- Thai, Vietnamese, Indonesian voiceover: on MAI's list, not Sume's tags
MAI-Voice-2.1 lists th-TH, vi-VN and id-ID; Sume's 16 voice tags do not. What that means, a 1-cent audition, and the Unicode trap in Vietnamese.
- Translate text in an image but keep brand names: the edit prompt
A prompt pattern for translating the words in an image on Sume while leaving brand names, prices and codes alone, with an Ideogram 4.5 request.
- Turkish TTS API: MAI-Voice-2.1 tr-TR vs Sume tr, and the dotted I bug
Turkish is on MAI-Voice-2.1 (tr-TR) and in Sume's voice tags (tr). Upper-casing a script in code turns i into I; the failure and a one-cent test to catch it.
Written by Sume