Shotstack tracks and clips vs Sume's audio spine and video slots
Shotstack layers clips on tracks; Sume Timeline 1.0 builds a video around one audio spine. How the edit models differ and when layered tracks win.

Shotstack builds an edit from layered tracks, and Sume builds a render from one audio track with video laid along it. If you need captions, overlays or several things on screen at once, Shotstack's model fits better. If your video is a voice-over with footage cut to it, Sume's is shorter to write.
Shotstack facts are from its edit fundamentals and API reference, read 2026-10-10. Sume facts are from Timeline 1.0 and Timeline compose.
The two edit models
A Shotstack edit has a timeline and an output. The timeline holds tracks, which are layers spanning its full length, with clips ordered top to bottom. A clip selects an asset (video, image, text or audio) and gives it a start time, a length, and optional transitions, effects and filters. The page warns that overlapping clips on one track cause flicker.
A Sume Timeline document has audio and video[]. The audio spine sets the length, and each slot has a start, a duration, an optional source_in and a fit of cover, contain, stretch or blur. There is a single video layer.
| Concept | Shotstack | Sume Timeline 1.0 |
|---|---|---|
| Top-level shape | timeline plus output | audio, video[], output |
| Layers | Many tracks | One video layer plus an optional soundtrack bed |
| Text and overlays | Text assets on a track | None in Timeline; use Timeline compose for a still over a video |
| What sets the length | The timeline's clips | audio.duration_seconds, 1 to 1800 |
| Auth header | x-api-key | Authorization: Bearer or x-api-key, never both |
| Submit and read | POST /render, then GET /render/{id}; status such as rendering, done | POST /v1/timeline-1.0/render, then GET /v1/jobs/:id/status; queued, processing, completed |
| Done notice | A callback URL in the edit | mode: "webhook" with webhook_url |
Why the audio-first model exists
Sume's design assumes narration is the spine. video[0].start must be 0, starts must increase, and coverage may stop at most 0.5 seconds before the end of the spine. You can duck a soundtrack bed under the spine with duck_db, and use up to 20 gapless audio.parts[] slices joined at sample level with no re-TTS. That is convenient for explainers and ads, and unhelpful for a montage with no voice.
Porting a simple Shotstack edit
A Shotstack edit with one video track maps to a Timeline. Import each file first with POST /v1/media-imports, because Timeline rejects off-host URLs. Put the narration in audio.url, set each clip's start and duration, and call POST /v1/timeline-1.0/plan for an unbilled check. Text tracks do not carry over. Rendering costs $0.10 per ceil output minute per the docs, with the live rate in GET /v1/catalog; Shotstack's pricing is outside this post because it was not read for it.
Where layered tracks win
Be plain about the gap. Shotstack's tracks let you put a logo on every second of a video, run a text lower-third over a clip, and stack a picture-in-picture without leaving the one JSON. Sume's Timeline has one video layer and a soundtrack bed. If your edit looks like a news segment, with several things on screen, Shotstack matches the shape of the work and Timeline does not.
Sume's model wins when the picture is meant to follow a voice. You set the narration length, drop clips under it, and the render service handles chunking past 12 segments with render.strategy set to auto. You also get POST /v1/timeline-1.0/plan as a free dry run, which is useful when an automated pipeline builds the document and you want to reject a bad one before it costs anything.
Sources
Related posts
More in Comparisons
- Spreadsheet rows to videos: Creatomate or Sume bulk runs?
Creatomate ties a template to a spreadsheet; Sume bulk runs queue up to 100 Format runs at concurrency 1 to 16. How rows map, and how many queues 250 rows take.
- Suno Italy probe: four terms clauses to check before ad music
Italy's AGCM opened a probe into Suno's terms on Oct 6, 2026. The four flagged clauses, and what to check in any AI music tool before putting a track in an ad.
- Suno v6-mini for all users vs Sume Music at $0.125 a track
Suno v6-mini is open to every user, while v6 and v6-wild need Pro or Premier. Sume does not list Suno; it sells one Lyria track for a flat $0.125.
- Synthesia API rate limits by tier vs Sume's queue and 429s
Synthesia caps writes at 60 to 120 a minute by tier and answers 429. Sume queues accepted jobs and returns queue_full or rate_limited. Read 2026-10-10.
Written by Sume