Teacher explainer avatar clips for a semester, with think-pause beats
Thirty-six 45-second avatar explainers for a 12-week term cost $298.08 on Sume's standard tier. A silence scene gives students a think-pause inside the clip.
Three 45-second avatar explainers a week for a 12-week term is 36 clips and 1,620 seconds of video, which comes to $298.08 on Sume's standard tier and $396.90 on plus at list rates with no product image. A multi-scene request can also include a silence scene, so the clip pauses for a question before the answer.
The prices are from Sume API pricing, and the scene format from Generate avatar video, read 2026-10-03. School policies on AI-made content vary, so check yours before publishing.
What does a term cost?
Each explainer is a single job in the 4 to 60 second window, so 45 seconds is comfortably inside it. Costs scale linearly with seconds, which makes a term easy to estimate: multiply clips by seconds by the tier rate.
| Tier | Per 45-second clip | 36 clips (1,620 s) |
|---|---|---|
| standard | $8.28 | $298.08 |
| plus | $11.025 | $396.90 |
| max | $24.75 | $891.00 |
How do you build a think-pause?
Use video_inputs instead of a single script. Each scene has a voice: a text voice with a script and a duration, or a silence voice with only a duration. A silence scene is a non-speaking beat, so the avatar stays on screen while students think. The total planned duration must still land in the 4 to 60 second window, and the current execution uses one avatar and one shared scene for the whole video.
For a 45-second explainer, a plan could be a 15-second setup of the problem, a 10-second silence for the class to work, and a 20-second worked answer. The scene durations add to 45 seconds, and this post treats the whole 45 seconds as the billed length to be safe, since the pricing page lists a rate per second without breaking out silence.
- Spoken scenes take exactly one of
scriptorinput_text. - A silence scene must not carry script text.
- Keep the background prompt identical across scenes, because the scenes resolve to one shared scene.
Captions and review
Inline captions are optional and burn into the final MP4 using the spoken text, which helps students watching without sound and anyone reading along. If caption styling fails, the avatar job can still succeed with a clean video and captions.status set to failed, so check that field before you publish.
Review the first frame of the first explainer with a preview, then render the rest at the tier you chose. Teachers who want a cheaper draft pass can preview at any tier, since preview stills are tier-independent and generate-video takes the final quality.
A cheaper path is to render the first three weeks at standard, watch how students use them, and only then commit the term. Nine 45-second clips cost $74.52 on standard, a quarter of the full term. If a topic needs a sharper render, such as one with fine diagram work behind the avatar, that single clip can be re-rendered at plus for $11.03 without touching the rest.
Sources
Related posts
More in Use cases
- Teams translated captions vanish after the meeting: keep them
Microsoft says Teams translated captions and transcripts are only available during the meeting. Caption the recording afterwards with Sume STT and burned cues.
- Template bulk edits on YouTube Shorts: what to vary per row
YouTube's Oct 1, 2026 originality update names template-based bulk changes as not original. How to make each row of a Sume bulk run differ in substance.
- Ten reference images for one product still: what goes in each slot
FLUX 3 Image takes up to ten references, GPT Image 2.5 on Sume takes sixteen, Ideogram 4.5 five. A slot plan and a script that trims it to each model's cap.
- Avatar reaction 3.49 vs 3.06 predicted: run your own two-clip test
HeyGen's survey says avatar users saw warmer reactions than skeptics predicted. Test your own audience with two Sume avatar clips, one script and safe retries.
Written by Sume