60-second explainer Short end to end: $7.84 with Wan 3.0 clips

A 60-second explainer built from TTS, twelve 5-second Wan 3.0 clips, one Timeline render and captions costs $7.84 on Sume. The clips are $7.50 of it.

5 min readSume
All posts

A 60-second explainer Short made from a voiceover, twelve 5-second Wan 3.0 clips, a Timeline render and captions comes to $7.84 on Sume, and the twelve clips are $7.50 of it. YouTube Shorts accept vertical videos up to 3 minutes (read 2026-10-09), so 60 seconds leaves room to extend.

The bill, line by line

Assumptions: a 900-character script (the TTS cost depends on characters, not on seconds), 720p clips, a 60-second audio spine, and one standalone caption job. Captions are priced for videos up to 60 seconds, which is why the explainer stops there.

60-second explainer, line items (read 2026-10-09)
StepCalculationCost
TTS voiceover900 chars x $0.0475 per 1,000$0.04275
Wan 3.0 720p clips12 x 5 s x $0.125/s = 12 x $0.625$7.50
Timeline 1.0 renderceil(60 s / 60) = 1 min x $0.10$0.10
Standalone captionsone job, video up to 60 s$0.20
Total$7.84 (7.84275 exact)

Assemble the render

After the twelve clips and the voice file exist as media.sume.com artifacts, one Timeline call lays them on the spine. Declare the audio duration you measured, and give every slot a start that increases from 0. Sume chunks the render automatically past 12 segments.

import os, requests

H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
     "Content-Type": "application/json", "Idempotency-Key": "explainer-001"}
clips = [f"https://media.sume.com/artifacts/artf_demo/clip{i}.mp4" for i in range(12)]
body = {
    "audio": {"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
              "duration_seconds": 60},
    "video": [{"source_url": u, "start": i * 5, "duration": 5,
               **({"transition": {"type": "fade", "duration": 0.25}} if i else {})}
              for i, u in enumerate(clips)],
}
r = requests.post("https://api.sume.com/v1/timeline-1.0/render",
                  headers=H, json=body, timeout=60)
r.raise_for_status()
print(r.json())

Limits and gotchas

  • Check the plan first: POST /v1/timeline-1.0/plan is unbilled and returns the estimated cost.
  • Wan 3.0 accepts 2 to 30 seconds per clip; fewer, longer clips are cheaper to assemble but change nothing in the per-second price.
  • Captions need audible speech on the clip; use your voice file as the spine.
  • Eight or more adjacent fades in a row are refused; insert a hard cut.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume