How many clips for a 10-minute faceless video? Timeline slot math
A 10-minute faceless video fits Timeline 1.0 comfortably: up to 200 slots, a render near $1.00, and limits on fades and single-pass renders to plan around.

A 10-minute faceless video on Sume's Timeline 1.0 can use anywhere from one to 200 video slots, so the practical answer is 60 to 120 clips at 5 to 10 seconds each. The render itself is $0.10 per output minute, about $1.00 for 10 minutes; the clips you generate or source cost extra.
What are the hard limits?
video[0].start must be 0, later starts must increase, and declared starts are authoritative. The compiler compensates for crossfades rather than shifting your starts.
| Limit | Value | Why it matters |
|---|---|---|
| audio.duration_seconds | 1 to 1800 | A 10-minute episode is 600 s, a third of the cap |
| video[] slots | 1 to 200 | 600 s over 200 slots averages 3 s each at the maximum |
| Slot duration | at least 0.2 s | Fast cuts are allowed |
| Coverage past the spine | at most 0.5 s | Slots must reach the end of the voiceover |
| Adjacent fades | 8 chained, then insert a hard cut | too_many_chained_transitions |
| render.strategy single | 12 slots or fewer | Above that use auto or chunked |
| Audio parts | 20 per spine | Join sliced voiceover without re-synthesis |
How many clips fit the pacing you want?
Divide the length by the shot length. At 6 seconds a shot, 600 seconds needs 100 slots. At 4 seconds it needs 150. At 3 seconds it hits the 200-slot ceiling. So the real constraint for a 10-minute video is how many shots you are willing to make or source, not the slot limit. A talking explainer with a new picture every sentence is plausible; a new picture every second is not.
What does a 10-minute episode cost end to end?
The render reserves ceil(audio.duration_seconds / 60) minutes, so 600 seconds is 10 minutes and $1.00 at the published rate. Check it first with POST /v1/timeline-1.0/plan, which is unbilled and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros. A plan cannot predict warnings about short sources being padded or looped.
What does the request look like?
Every URL must already be this workspace's media.sume.com media. A shortened body for two slots; a real episode has the full list. Default mode is async, so poll GET /v1/jobs/:id/status or use a webhook.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: faceless-ep7" \
-d '{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 600 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/s1.mp4", "start": 0, "duration": 6 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/s2.mp4", "start": 6, "duration": 6 }
]
}'What goes wrong on long episodes?
The omitted output.fps matches the sources, and a rate that differs from a source's repeats or drops frames and reports output_fps_resamples_sources. Mixed-rate B-roll across 100 clips is a common source of that warning, so set one rate. Short sources are padded or looped with soft warnings, so read warnings[] in the result, which is a timeline_render with video_url, duration_seconds, segment_count and billable_minutes.
Default output is 1080 by 1920. For a landscape channel set output.width and output.height yourself, and use fit: "blur" where a vertical clip lands in a wide frame.
Should you render in one pass or in chunks?
Leave render.strategy at auto. It chunks once a render passes 12 segments, which is the case for almost every 10-minute episode. Asking for single above 12 slots is refused as render_strategy_unsafe. If you want finer control over reruns, you can build the episode as three 3 to 4-minute renders and join them, but a single render keeps the voiceover continuous and is simpler to reason about.
Keep a spreadsheet of slot, start, duration and clip URL. When a clip needs replacing you can edit one row and re-run the plan call before paying for another render.
What is a sensible starting plan?
For a first episode, aim for about 80 slots at 7 to 8 seconds, one per spoken idea. Run the plan call, read segment_count and estimated_cost_usd_micros, render, and watch the result for warnings[]. If the pacing feels slow, split the longest slots; if you hit the fade limit, replace some fades with hard cuts. Plan, render and review is faster than guessing the perfect count up front.
What does Sume not do here?
Timeline 1.0 assembles; it does not write the script, pick clips, or voice the narration. Voiceover must already be a Sume-hosted audio file. For the generation side, see the video generation docs.
Sources
Related posts
More in Use cases
- Instagram Reels AI translation: Korean and Japanese added July 2026
Meta's July 14, 2026 update adds French, German, Italian, Japanese and Korean to Reels translation on Instagram. Here is the full language list and the dates.
- Korean captions on product clips: use korean-ad, not slam
slam, punch and tiktok-green have no Hangul glyphs and return 400 on Korean copy. Use korean-ad with language ko for Korean product clips; $0.20 a job.
- Find product moments in a live-commerce replay and trim them out
Run one transcript on a live-selling replay, find the sentences where each product is pitched, then cut each range with video-trim at $0.02 a clip.
- Live replay highlight reel in one render: detach once, no trims
Detach a live replay's audio once ($0.01), then build a 30-second highlight reel in one Timeline render using audio parts and source_in. No per-clip trim jobs.
Written by Sume