Explainer video from a script: Sonic 3.6 voice and Timeline

Build a narrated explainer by generating the voiceover with the TTS Router on sonic-3.6, then rendering slots on Timeline 1.0 at $0.10 per output minute.

5 min readSume
All posts

An explainer video is a voiceover plus pictures that follow it. On Sume you generate the narration with POST /v1/tts-router/generate on sonic-3.6, import it and your clips, then call POST /v1/timeline-1.0/render with the audio as the spine and ordered video[] slots. Timeline bills $0.10 per ceil output minute, and it runs ffmpeg, not a model.

Voice first, because it sets the length

In an explainer the audio decides everything. Each scene lasts as long as its sentence. So write the script, generate the voice, and read the duration before you plan visuals. The TTS Router takes a required model from its catalog, either transcript or a transcript_source, and one voice selector such as avatar_handle or voice.id.

The router seeds Cartesia Sonic only: sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview. The alias sonic-latest resolves to sonic-3.6. Pin the explicit id so a future alias change does not move your sound.

The Timeline document

Timeline 1.0 takes one audio spine and an ordered list of video slots, and returns one MP4. audio.duration_seconds runs from 1 to 1800, and video[] holds 1 to 200 slots. Every URL must already be an artifact or asset in your workspace on media.sume.com, so import files first with POST /v1/media-imports. An Idempotency-Key is required.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: explainer-001" \
  -d '{
    "audio": {"url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 24},
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4", "start": 8, "duration": 16,
       "transition": {"type": "fade", "duration": 0.25}}
    ]
  }'

Preflight for free

Plan before you pay for a render. A mistake in a slot, such as a URL outside your workspace, shows up in the plan response and costs nothing.

POST /v1/timeline-1.0/plan runs the same schema checks and compiler and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. It creates no job and reserves no credits, and it needs no idempotency key. A plan cannot predict warnings about short sources that get padded or looped, so still read warnings[] on the real result.

Matching visuals to sentences

Split the script at sentence boundaries and give each sentence one slot. A slot is a source_url, a start and a duration, so the length of each slot is the length of the spoken line. Put a short fade between topics with the transition field, and keep a hard cut inside one idea.

If a source clip is shorter than its slot, Timeline pads or loops it and reports a soft warning, not a failure. Read the warnings, then regenerate the clip at the right length rather than letting it loop on screen.

What an explainer costs

The public Timeline rate is $0.10 per ceil output minute, and the reserve is ceil(audio.duration_seconds / 60) minutes. A 90-second explainer therefore reserves 2 minutes, which is $0.20 for assembly. The voice and the clips are billed separately by their own rates in the catalog.

Timeline rate from the Sume Timeline docs, read 2026-10-05
Audio lengthReserved minutesAssembly cost
24 seconds1$0.10
90 seconds2$0.20
5 minutes5$0.50

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume