How many clips for a 10-minute faceless video? Timeline slot math

A 10-minute faceless video fits Timeline 1.0 comfortably: up to 200 slots, a render near $1.00, and limits on fades and single-pass renders to plan around.

5 min readSume
All posts

A 10-minute faceless video on Sume's Timeline 1.0 can use anywhere from one to 200 video slots, so the practical answer is 60 to 120 clips at 5 to 10 seconds each. The render itself is $0.10 per output minute, about $1.00 for 10 minutes; the clips you generate or source cost extra.

What are the hard limits?

video[0].start must be 0, later starts must increase, and declared starts are authoritative. The compiler compensates for crossfades rather than shifting your starts.

Timeline 1.0 limits that shape a faceless episode (read 2026-10-02)
LimitValueWhy it matters
audio.duration_seconds1 to 1800A 10-minute episode is 600 s, a third of the cap
video[] slots1 to 200600 s over 200 slots averages 3 s each at the maximum
Slot durationat least 0.2 sFast cuts are allowed
Coverage past the spineat most 0.5 sSlots must reach the end of the voiceover
Adjacent fades8 chained, then insert a hard cuttoo_many_chained_transitions
render.strategy single12 slots or fewerAbove that use auto or chunked
Audio parts20 per spineJoin sliced voiceover without re-synthesis

How many clips fit the pacing you want?

Divide the length by the shot length. At 6 seconds a shot, 600 seconds needs 100 slots. At 4 seconds it needs 150. At 3 seconds it hits the 200-slot ceiling. So the real constraint for a 10-minute video is how many shots you are willing to make or source, not the slot limit. A talking explainer with a new picture every sentence is plausible; a new picture every second is not.

What does a 10-minute episode cost end to end?

The render reserves ceil(audio.duration_seconds / 60) minutes, so 600 seconds is 10 minutes and $1.00 at the published rate. Check it first with POST /v1/timeline-1.0/plan, which is unbilled and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. A plan cannot predict warnings about short sources being padded or looped.

What does the request look like?

Every URL must already be this workspace's media.sume.com media. A shortened body for two slots; a real episode has the full list. Default mode is async, so poll GET /v1/jobs/:id/status or use a webhook.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: faceless-ep7" \
  -d '{
    "audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 600 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/s1.mp4", "start": 0, "duration": 6 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/s2.mp4", "start": 6, "duration": 6 }
    ]
  }'

What goes wrong on long episodes?

The omitted output.fps matches the sources, and a rate that differs from a source's repeats or drops frames and reports output_fps_resamples_sources. Mixed-rate B-roll across 100 clips is a common source of that warning, so set one rate. Short sources are padded or looped with soft warnings, so read warnings[] in the result, which is a timeline_render with video_url, duration_seconds, segment_count and billable_minutes.

Default output is 1080 by 1920. For a landscape channel set output.width and output.height yourself, and use fit: "blur" where a vertical clip lands in a wide frame.

Should you render in one pass or in chunks?

Leave render.strategy at auto. It chunks once a render passes 12 segments, which is the case for almost every 10-minute episode. Asking for single above 12 slots is refused as render_strategy_unsafe. If you want finer control over reruns, you can build the episode as three 3 to 4-minute renders and join them, but a single render keeps the voiceover continuous and is simpler to reason about.

Keep a spreadsheet of slot, start, duration and clip URL. When a clip needs replacing you can edit one row and re-run the plan call before paying for another render.

What is a sensible starting plan?

For a first episode, aim for about 80 slots at 7 to 8 seconds, one per spoken idea. Run the plan call, read segment_count and estimated_cost_usd_micros, render, and watch the result for warnings[]. If the pacing feels slow, split the longest slots; if you hit the fade limit, replace some fades with hard cuts. Plan, render and review is faster than guessing the perfect count up front.

What does Sume not do here?

Timeline 1.0 assembles; it does not write the script, pick clips, or voice the narration. Voiceover must already be a Sume-hosted audio file. For the generation side, see the video generation docs.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume