Ten micro-lessons from ten scripts: plan first, $5.29 to render

A ten-lesson micro-course on Sume costs $5.29: per lesson 600 characters of TTS, two stills, a render and cues. Free /plan checks first, read 2026-10-10.

4 min readSume
All posts

Ten 45-second micro-lessons cost about $5.29 on Sume, or $0.5285 each: 600 characters of TTS ($0.0285), two stills ($0.20), one Timeline render ($0.10) and one cue caption job ($0.20). The Timeline plan call is free, so each lesson can be checked before it is rendered.

A micro-course is a production line. The lesson shape never changes: a title still, a diagram-like still, a voice and a key-line caption. That makes the cost predictable and the failures cheap. Rates are from the Sume pricing code (read 2026-10-10).

What is the endpoint chain per lesson?

Each lesson is its own set of jobs with its own idempotency keys, for example lesson-04-tts.

  • POST /v1/tts-1.0/generate with the 600-character script, timestamps.words: true.
  • POST /v1/images x2 with google/nano-banana-2.1: a scene still and a concept still; no text in the pictures.
  • POST /v1/timeline-1.0/plan with the real request: free, returns validation and output length.
  • POST /v1/timeline-1.0/render with the same body, then POST /v1/video-captions with cues for the one key line.

What does the course cost?

Everything except the voice is flat per lesson.

Micro-course cost by lesson count (read 2026-10-10)
LessonsTTSStillsRenderCaptionsTotal
5$0.143$1.00$0.50$1.00$2.64
10$0.285$2.00$1.00$2.00$5.29
20$0.570$4.00$2.00$4.00$10.57

What are the limits?

Each render is billed per started output minute, so keep every lesson under 60 seconds to pay for one. A 600-character script is roughly 45 seconds at a default pace, but test one first: generation_config.speed ranges 0.6 to 1.5 if a lesson runs over. A failed /plan costs nothing and names the slot or field to fix.

Run lessons together, but within your workspace's concurrency; ten at once may queue. A callback per job is cleaner than ten polling loops. See Jobs and results and Webhooks, and read the 7-minute explainer for plan-first detail.

Do not teach anything where the pictures must be exact, like a medical figure. The model makes plausible images, not diagrams.

How do I keep the series consistent?

Use the same image style words in all prompts, the same caption style and the same voice. Keep the script template fixed so that cost stays at the same number for each lesson. If the series needs a presenter, an English-only avatar is possible; see AI avatar for online course videos, but its seconds cost far more than voice-over.

Pre-publish checklist

Render lesson one alone and watch it with a learner in mind. A shared structure means one fix applies to all ten.

Write each script to fit the same word count, so every render stays one started minute.

  • Run /plan before every render.
  • Use one caption style for the whole series.
  • Keep a table of lesson, job ids and status in your own records.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume