Stitch AI clips into one MP4 with Sume Timeline in Python
POST /v1/timeline-1.0/render with audio mode silence joins clips into one MP4 for $0.10 a minute. A short Python script, with the plan call first.

To stitch AI clips into one MP4 with Sume, call POST /v1/timeline-1.0/render with audio.mode set to silence, an audio.duration_seconds equal to the total length, and one video[] slot per clip. The render costs $0.10 per ceil(output minute) and adds no model inference.
The script below plans first, which is free, and then renders. Clips must already be media.sume.com files in your workspace.
The script
Silence mode declares the length with no audio file. In that mode you must not send url, parts, gain_db or source_in. The first slot starts at 0, later starts increase, and a transition appears only on slots after the first.
import os
import requests
BASE = "https://api.sume.com/v1/timeline-1.0"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
M = "https://media.sume.com/artifacts/artf_demo/"
body = {
"audio": {"mode": "silence", "duration_seconds": 15},
"video": [
{"source_url": M + "a.mp4", "start": 0, "duration": 5},
{"source_url": M + "b.mp4", "start": 5, "duration": 5,
"transition": {"type": "fade", "duration": 0.25}},
{"source_url": M + "c.mp4", "start": 10, "duration": 5},
],
"output": {"fps": 24},
}
plan = requests.post(f"{BASE}/plan", json=body, headers=H, timeout=30)
print(plan.json())
r = requests.post(f"{BASE}/render", json=body, timeout=30,
headers={**H, "Idempotency-Key": "stitch-demo-001"})
print(r.status_code, r.json())What the response gives you
A successful submit returns a job with type: timeline_render and model: sume/timeline-1.0. Poll GET /v1/jobs/:id/status, then read GET /v1/jobs/:id/result when it is result_ready. The result has kind: timeline_render, video_url, duration_seconds, segment_count, billable_minutes and any warnings[]. Soft warnings, such as a padded or looped short source, are not failures.
To wait for a finished job in the response, send mode: sync. Sume waits up to 30 seconds and returns 200, or returns 202 if the job is not done.
The rules that trip people up
| Rule | Value |
|---|---|
| Transition types | fade, wipeleft, wiperight, slideup, slidedown, dissolve |
| Transition length | up to 1 s and up to 50% of the shorter neighbor |
| Fit | cover (default), contain, stretch, blur |
| Default output size | 1080x1920 MP4 |
| output.fps | 24, 25, 30 or 60; omit to match the sources |
| Coverage | slots may stop at most 0.5 s before the end of the spine |
Pick the fps on purpose
If you omit output.fps, the render follows the sources. If you set a rate that differs from a source, the job repeats or drops frames and warns output_fps_resamples_sources, which shows as judder on motion. AI clips from different models often have different rates, so set one rate for the whole cut and conform each clip first if needed with the output field of video trim.
If you need sound, replace the silence spine with a real audio file or add a soundtrack bed. The Timeline docs list the fields.
Common refusals and their fixes
A render that starts at a non-zero time is refused with timeline_must_start_at_zero. A fade on the first slot is transition_on_first_segment. Slots that overlap beyond the crossfade are segment_overlap. A slot list that stops more than 0.5 seconds before the end of the spine fails the coverage rule. A hand-written filtergraph, codec or crf key is rejected with a 400, because the server compiles ffmpeg itself.
Using a public URL that is not on media.sume.com is unsupported_media_source; import first with POST /v1/media-imports.
Cost of the example
The 15-second render bills ceil(15/60) = 1 minute, so $0.10. The plan call is free. If the three clips are 5-second jobs on a model that charges R dollars per clip, the full cut costs 3R + $0.10. The render adds no model inference.
Next steps after the render
Download or reuse the video_url from the result. If the cut needs captions or a logo, add them in a separate step and render again, since Timeline joins and fades but does not draw text. For a half-banner with a still above a clip, use timeline compose first, then place its MP4 in a slot.
Before you build
Before you build, read the linked Sume docs page for the exact request fields, limits and prices, because those pages are the source of truth and can change. Run one short, cheap test with your own material first, check the output in a player and in your editor, and only then scale to the full shot list. Keep every job id and file you approve, so a later change never forces you to regenerate work that was already signed off. Note that this post describes Sume's catalog and tools; Sume does not run Luma Ray 3.2, and nothing here claims HDR or EXR output.
Sources
Related posts
More in Developers
- Stitch three STT chunks: add each range start to word times
Sume STT times count from each chunk's own start, so add the chunk's range start to every word and sentence. A short Python function that does it.
- Stop an Omni draft batch at $5: sum usage.cost from each poll
A Python loop that submits 360p Omni drafts one at a time, adds the Sume usage.cost of each finished job, and stops before the next one would cross a $5 cap.
- Can Strands Decider 2B pick which Sume MCP tool to call?
Strands Decider 2B scores options you give it. Build them from Sume's tools_list, then check the pick with tools_schema and dry_run before any paid call.
- Stripe allows 16 webhook endpoints; Sume takes a webhook URL per job
Stripe registers up to 16 endpoints. Sume has no registry: each job or Format run item carries its own public HTTPS webhook_url, signed with one secret.
Written by Sume