60-second avatar explainer: one $15.48 job or two 30-second jobs
A 60-second Sume Avatar 1.0 explainer with a product image is $15.48 on plus, $11.64 standard, $34.80 max. One job or two 30-second jobs, and what changes.
A 60-second product explainer on Sume Avatar 1.0 with a product image costs $11.64 on standard, $15.48 on plus and $34.80 on max, which is the same whether you run one 60-second job or two 30-second jobs, to within chunk rounding. Run one job if the explainer is one continuous take; split it if you want to retry the halves separately or change the scene between them.
Sixty seconds is the top of the window: the Avatar Video docs (read 2026-10-06) accept scripts and multi-scene plans that Sume estimates at 4 to 60 seconds inclusive. At 2.8 words a second that is about 168 words.
The numbers
The rate is per second, so splitting does not change the list price. Sume reserves the cost at submit from chunk-planned billable seconds, so read the estimate on a preview rather than trusting a hand calculation to the cent.
| Tier | Per second with product | One 60 s job | Two 30 s jobs |
|---|---|---|---|
| standard | $0.194 | $11.64 | $5.82 each, $11.64 |
| plus (default) | $0.258 | $15.48 | $7.74 each, $15.48 |
| max | $0.58 | $34.80 | $17.40 each, $34.80 |
One job or two
One job keeps a single avatar, scene and pose through the whole explainer, and a multi-scene video_inputs plan lets you put a hook, a demo and a call to action in one render, including silence beats. The current execution supports one resolved avatar per final video and expects the scene backgrounds to resolve to one shared scene.
Two jobs cost the same but let you re-render only the half that went wrong. The tradeoff is the join: you add a Timeline 1.0 render at $0.10 per started output minute, which is $0.10 for a 60-second result, and the two halves will not share an unbroken take. Cut on a scene change in the script so the seam reads as an edit.
A practical rule
Write the explainer as one script, split it only where the story has a natural beat, and render the first half on standard to see how the avatar delivers it. If the first half is right, render the second. If you are unsure of the product image, preview the first frame before either job. A wrong script in one 60-second job is the most expensive mistake in this table, so keep the script review before the render, not after it.
Sources
Related posts
More in Sume Avatar 1.0
- AI clone from 2 minutes of video or one photo: what each needs
Tavus asks for two minutes of 1080p video and written consent. Sume Avatar 1.0 starts from a prompt, traits or one photo URL. The inputs compared, dated.
- Avatar clip under 4 seconds: add a silence beat or a longer line
Sume avatar scripts must plan to 4 seconds or more. For a one-liner, add a silence beat in video_inputs or lengthen the line instead of padding with filler.
- Talking avatar from an approved TTS file: image-to-video audio_url
To animate an avatar from a voiceover you already approved, send its Sume-hosted audio_url and duration_seconds to Avatar 1.0 image-to-video. Limits inside.
- AI avatar without a photo: Sume props input (age, sex, ethnicity)
No photo of your presenter? Sume creates an avatar from a text prompt or structured props such as age, sex and ethnicity. One POST, one flat creation price.
Written by Sume