Shotstack tracks and clips vs Sume's audio spine and video slots

Shotstack layers clips on tracks; Sume Timeline 1.0 builds a video around one audio spine. How the edit models differ and when layered tracks win.

5 min readSume
All posts

Shotstack builds an edit from layered tracks, and Sume builds a render from one audio track with video laid along it. If you need captions, overlays or several things on screen at once, Shotstack's model fits better. If your video is a voice-over with footage cut to it, Sume's is shorter to write.

Shotstack facts are from its edit fundamentals and API reference, read 2026-10-10. Sume facts are from Timeline 1.0 and Timeline compose.

The two edit models

A Shotstack edit has a timeline and an output. The timeline holds tracks, which are layers spanning its full length, with clips ordered top to bottom. A clip selects an asset (video, image, text or audio) and gives it a start time, a length, and optional transitions, effects and filters. The page warns that overlapping clips on one track cause flicker.

A Sume Timeline document has audio and video[]. The audio spine sets the length, and each slot has a start, a duration, an optional source_in and a fit of cover, contain, stretch or blur. There is a single video layer.

Shotstack facts from its docs (read 2026-10-10); Sume facts from the Timeline 1.0 docs.
ConceptShotstackSume Timeline 1.0
Top-level shapetimeline plus outputaudio, video[], output
LayersMany tracksOne video layer plus an optional soundtrack bed
Text and overlaysText assets on a trackNone in Timeline; use Timeline compose for a still over a video
What sets the lengthThe timeline's clipsaudio.duration_seconds, 1 to 1800
Auth headerx-api-keyAuthorization: Bearer or x-api-key, never both
Submit and readPOST /render, then GET /render/{id}; status such as rendering, donePOST /v1/timeline-1.0/render, then GET /v1/jobs/:id/status; queued, processing, completed
Done noticeA callback URL in the editmode: "webhook" with webhook_url

Why the audio-first model exists

Sume's design assumes narration is the spine. video[0].start must be 0, starts must increase, and coverage may stop at most 0.5 seconds before the end of the spine. You can duck a soundtrack bed under the spine with duck_db, and use up to 20 gapless audio.parts[] slices joined at sample level with no re-TTS. That is convenient for explainers and ads, and unhelpful for a montage with no voice.

Porting a simple Shotstack edit

A Shotstack edit with one video track maps to a Timeline. Import each file first with POST /v1/media-imports, because Timeline rejects off-host URLs. Put the narration in audio.url, set each clip's start and duration, and call POST /v1/timeline-1.0/plan for an unbilled check. Text tracks do not carry over. Rendering costs $0.10 per ceil output minute per the docs, with the live rate in GET /v1/catalog; Shotstack's pricing is outside this post because it was not read for it.

Where layered tracks win

Be plain about the gap. Shotstack's tracks let you put a logo on every second of a video, run a text lower-third over a clip, and stack a picture-in-picture without leaving the one JSON. Sume's Timeline has one video layer and a soundtrack bed. If your edit looks like a news segment, with several things on screen, Shotstack matches the shape of the work and Timeline does not.

Sume's model wins when the picture is meant to follow a voice. You set the narration length, drop clips under it, and the render service handles chunking past 12 segments with render.strategy set to auto. You also get POST /v1/timeline-1.0/plan as a free dry run, which is useful when an automated pipeline builds the document and you want to reject a bad one before it costs anything.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume