Ten micro-lessons from ten scripts: plan first, $5.29 to render
A ten-lesson micro-course on Sume costs $5.29: per lesson 600 characters of TTS, two stills, a render and cues. Free /plan checks first, read 2026-10-10.

Ten 45-second micro-lessons cost about $5.29 on Sume, or $0.5285 each: 600 characters of TTS ($0.0285), two stills ($0.20), one Timeline render ($0.10) and one cue caption job ($0.20). The Timeline plan call is free, so each lesson can be checked before it is rendered.
A micro-course is a production line. The lesson shape never changes: a title still, a diagram-like still, a voice and a key-line caption. That makes the cost predictable and the failures cheap. Rates are from the Sume pricing code (read 2026-10-10).
What is the endpoint chain per lesson?
Each lesson is its own set of jobs with its own idempotency keys, for example lesson-04-tts.
POST /v1/tts-1.0/generatewith the 600-character script,timestamps.words: true.POST /v1/imagesx2 withgoogle/nano-banana-2.1: a scene still and a concept still; no text in the pictures.POST /v1/timeline-1.0/planwith the real request: free, returns validation and output length.POST /v1/timeline-1.0/renderwith the same body, thenPOST /v1/video-captionswithcuesfor the one key line.
What does the course cost?
Everything except the voice is flat per lesson.
| Lessons | TTS | Stills | Render | Captions | Total |
|---|---|---|---|---|---|
| 5 | $0.143 | $1.00 | $0.50 | $1.00 | $2.64 |
| 10 | $0.285 | $2.00 | $1.00 | $2.00 | $5.29 |
| 20 | $0.570 | $4.00 | $2.00 | $4.00 | $10.57 |
What are the limits?
Each render is billed per started output minute, so keep every lesson under 60 seconds to pay for one. A 600-character script is roughly 45 seconds at a default pace, but test one first: generation_config.speed ranges 0.6 to 1.5 if a lesson runs over. A failed /plan costs nothing and names the slot or field to fix.
Run lessons together, but within your workspace's concurrency; ten at once may queue. A callback per job is cleaner than ten polling loops. See Jobs and results and Webhooks, and read the 7-minute explainer for plan-first detail.
Do not teach anything where the pictures must be exact, like a medical figure. The model makes plausible images, not diagrams.
How do I keep the series consistent?
Use the same image style words in all prompts, the same caption style and the same voice. Keep the script template fixed so that cost stays at the same number for each lesson. If the series needs a presenter, an English-only avatar is possible; see AI avatar for online course videos, but its seconds cost far more than voice-over.
Pre-publish checklist
Render lesson one alone and watch it with a learner in mind. A shared structure means one fix applies to all ten.
Write each script to fit the same word count, so every render stays one started minute.
- Run
/planbefore every render. - Use one caption style for the whole series.
- Keep a table of lesson, job ids and status in your own records.
Sources
Related posts
More in Use cases
- Thrift and resale listing videos from one flat-lay photo
A reseller can turn a flat-lay photo into a 6-second listing clip with Wan 3.0 first-frame video for $0.38 at 480p. Rules for honest motion and a batch run.
- TikTok Spark Ads: 10,000 per account cap, un-authorize before delete
TikTok's Spark Ads page caps an Ads Manager account at 10,000 Spark Ads and says to un-authorize a video before deleting it. A rule for batches of Sume renders.
- Tire shop winter swap promo: voiceover plus timeline for about $1.27
A tire shop can build a 45-second winter-swap promo from three short clips, a text-to-speech voiceover and one timeline render. Itemized to the cent.
- Toast Websites video: 50 MB, landscape only, uploaded by hand
Toast Websites takes a brand video up to 50 MB in landscape only and does not pull it automatically. Make a 16:9 holiday-menu clip with Sume and trim it to fit.
Written by Sume