Avatar Presenter Training Video Over 60 Seconds: Split Into Jobs
A single avatar video job covers 4 to 60 seconds. Plan a 3-minute presenter video as three jobs and stitch them: costs at standard, plus and max quality.
The short answer
Sume accepts avatar video scripts that it estimates at 4 to 60 seconds, so a 3-minute presenter video is three jobs of up to 60 seconds, joined in Timeline 1.0. At standard quality ($0.184/s) 180 seconds is $33.12, at plus ($0.245/s) $44.10 and at max ($0.55/s) $99.00, plus one timeline render of 3 minutes ($0.30) and a one-time $0.95 to create the avatar.
How to split
Split the script at topic breaks, not at a word count, so each job ends on a complete thought. Keep an eye on the estimate: docs say Sume accepts 4 to 60 seconds inclusive, and a script that estimates longer must be made shorter or split. Give the three jobs the same avatar_handle, aspect_ratio (default 9:16) and quality, so the person looks the same across the cuts. Resolution is 720p.
For a multi-scene job, use ordered video_inputs. A voice.type: silence beat with a duration is a pause with no speech; use one between two points that need breathing room. The current execution supports one resolved avatar for each final video and expects the scene backgrounds to resolve to one shared scene.
| Line | Rate | Quantity | Total |
|---|---|---|---|
| Avatar creation (once) | $0.95 | 1 | $0.95 |
| Avatar video, standard | $0.184 per second | 180 s | $33.12 |
| Avatar video, plus (default) | $0.245 per second | 180 s | $44.10 |
| Avatar video, max | $0.55 per second | 180 s | $99.00 |
| Timeline 1.0 render | $0.10 per minute | 3 | $0.30 |
Review before the full run
Use an avatar video preview to see first-frame stills before you pay for a render, and generate the video from the preview id once you approve the framing. Stills only show the first frame, so they do not prove lip movement or pacing; watch the first job in full before you commit the other two.
Inline captions on the avatar job are an add-on to the estimate and are burned after generation; a failed caption stage is a soft failure and leaves a clean video_url. For three jobs, caption once on the stitched file instead if the total is under 60 seconds of speech per caption job.
One more budget line: a failed or rejected take. At plus quality, a full 60-second retry is $14.70. If you plan for one retry per three jobs, add that to the 180-second line, and consider standard quality for the first read-through, since 60 seconds at standard is $11.04.
- Use only an avatar from a person who consented to it.
- Disclose that the presenter is synthetic where your company policy or local rules ask for it.
- Pick
plusfor the first job, then decide whether the other two needmax.
Sources
Related posts
More in Use cases
- Face Swap a 15-Second Clip With a Consented Avatar: What It Costs
Beta Avatar Face Swap on Sume applies a ready avatar to a 4 to 15 second source video. Price at standard, plus and max, the avatar fee and the consent rules.
- Burn captions on a Claude Motion MP4 with cues, not speech-to-text
A Claude Motion export is a public-URL problem and a silent-clip problem. Sume video captions take authored cues, skip STT, and cost $0.20 per job up to 60 s.
- Can a 4:5 video be a YouTube Short? The page says square or vertical
YouTube's Shorts page says square or vertical and up to 3 minutes, and lists no ratio chart. A 4:5 feed render is taller than wide; how to make one on Sume.
- Caption once after stitching: four 15 s clips cost $0.30, not $0.90
The caption job is $0.20 per accepted job up to 60 s. Stitch four 15 s clips first ($0.10), then caption once ($0.20): $0.30, not $0.90.
Written by Sume