Explainer video from a script: Sonic 3.6 voice and Timeline
Build a narrated explainer by generating the voiceover with the TTS Router on sonic-3.6, then rendering slots on Timeline 1.0 at $0.10 per output minute.

An explainer video is a voiceover plus pictures that follow it. On Sume you generate the narration with POST /v1/tts-router/generate on sonic-3.6, import it and your clips, then call POST /v1/timeline-1.0/render with the audio as the spine and ordered video[] slots. Timeline bills $0.10 per ceil output minute, and it runs ffmpeg, not a model.
Voice first, because it sets the length
In an explainer the audio decides everything. Each scene lasts as long as its sentence. So write the script, generate the voice, and read the duration before you plan visuals. The TTS Router takes a required model from its catalog, either transcript or a transcript_source, and one voice selector such as avatar_handle or voice.id.
The router seeds Cartesia Sonic only: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The alias sonic-latest resolves to sonic-3.6. Pin the explicit id so a future alias change does not move your sound.
The Timeline document
Timeline 1.0 takes one audio spine and an ordered list of video slots, and returns one MP4. audio.duration_seconds runs from 1 to 1800, and video[] holds 1 to 200 slots. Every URL must already be an artifact or asset in your workspace on media.sume.com, so import files first with POST /v1/media-imports. An Idempotency-Key is required.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: explainer-001" \
-d '{
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
{"source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16,
"transition": {"type": "fade", "duration": 0.25}}
]
}'Preflight for free
Plan before you pay for a render. A mistake in a slot, such as a URL outside your workspace, shows up in the plan response and costs nothing.
POST /v1/timeline-1.0/plan runs the same schema checks and compiler and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. It creates no job and reserves no credits, and it needs no idempotency key. A plan cannot predict warnings about short sources that get padded or looped, so still read warnings[] on the real result.
Matching visuals to sentences
Split the script at sentence boundaries and give each sentence one slot. A slot is a source_url, a start and a duration, so the length of each slot is the length of the spoken line. Put a short fade between topics with the transition field, and keep a hard cut inside one idea.
If a source clip is shorter than its slot, Timeline pads or loops it and reports a soft warning, not a failure. Read the warnings, then regenerate the clip at the right length rather than letting it loop on screen.
What an explainer costs
The public Timeline rate is $0.10 per ceil output minute, and the reserve is ceil(audio.duration_seconds / 60) minutes. A 90-second explainer therefore reserves 2 minutes, which is $0.20 for assembly. The voice and the clips are billed separately by their own rates in the catalog.
| Audio length | Reserved minutes | Assembly cost |
|---|---|---|
| 24 seconds | 1 | $0.10 |
| 90 seconds | 2 | $0.20 |
| 5 minutes | 5 | $0.50 |
Sources
Related posts
More in Use cases
- Farm stand weekend reel: GPT Image 2.5 and Wan 3.0 clips for $1.41
Four 9:16 first frames from GPT Image 2.5 and four 5-second Wan 3.0 clips at 480p cost $1.41 with a Timeline render on Sume. The steps and the checks.
- Fashion lookbook video in 3:4 or 9:16: which Sume models take it
Wan 3.0, Seedance and MiniMax H3 take 3:4 and 9:16 on Sume; Kling and Omni do not take 3:4. Wan at 720p is $0.125 a second, H3 at 768p is $0.075.
- Feedback request video for an NPS survey: AI avatar script, cost
A 10-second avatar clip that asks for survey feedback: what to say, what to leave to the email, and the cost for 1,000 sends if you render per segment.
- Festive caption colors for holiday ads: slam style design overrides
TikTok's holiday guide asks for bold captions and festive overlays at peak. One caption job with design colors burns a red-and-green slam line for $0.20.
Written by Sume