Syren mixes the first 24 audio tracks; Sume's timeline takes 20 parts
Syren mixes only the first 24 audio tracks in an export. Sume's timeline takes one spine in up to 20 gapless parts plus one bed. What fits where.

Syren's help article says it mixes the first 24 audio tracks into your exported video; anything past 24 is not mixed (read 2026-10-11). Sume's Timeline 1.0 has a different model: one audio spine, which can be assembled from up to 20 gapless parts, plus one optional soundtrack bed underneath.
The Syren fact comes from How do I create a video with Syren?, read 2026-10-11. The Sume facts come from the Timeline 1.0 and Timeline audio pages. This is a limits comparison, not a quality ranking.
Two different audio models
Syren is a layered mixer: narration, music and effects are separate tracks that get summed, and Synthesia documents a cap on how many are included. Sume's timeline is a sequence: one spine decides the length of the output and the video slots are laid against it by start time. Layering exists only in one place, the soundtrack bed, which can loop, fade out for up to 10 seconds and duck under speech by 0 to 20 dB when the spine has real audio.
What each limit means in practice
For a talking presenter with a bed of music, both systems fit easily. The limit only bites when you assemble many voices or effects.
Two more timeline details help when you plan. The audio length you declare is the output length, and video slots may stop at most half a second before the end of the spine. Stills are static holds in the timeline, so a slide-style cutaway is a still with a duration rather than an animation.
| Question | Syren | Sume Timeline 1.0 |
|---|---|---|
| Tracks in the final mix | First 24 audio tracks | One spine plus one soundtrack bed |
| Pieces in the spine | Not stated | Up to 20 gapless parts |
| Output length | Up to 3 minutes | 1 to 1,800 seconds of audio |
| Music under speech | A track among the 24 | soundtrack with duck_db 0-20 and fade_out up to 10 s |
| Silent stretch | Not stated | audio.mode silence for a declared length |
When you would hit Sume's limit
A 20-part spine covers a 20-segment voice-over. If you have more than 20 spoken pieces, merge some of them first with the separate timeline audio job, which produces a reusable merged file at a fixed price per job (the catalog lists it at $0.01), then use that file as the spine. Because the join is sample-domain with no re-TTS, the voice does not change at the seams.
Sound effects are the real gap. Sume's timeline does not take a stack of effect tracks, so if your video needs many layered sounds, build the effect bed outside the timeline and feed it in as the single soundtrack. If the layered mix is the point of your project, Syren's 24-track model is closer to what you want.
A talking-head example
Take a 90-second founder update. On Sume it is two avatar jobs of 45 seconds each, both spoken by the avatar, joined by detaching each clip's audio into parts and rendering one timeline with an optional music bed ducked under the voice. That is two parts out of 20 and one bed, nowhere near a limit, and the extra timeline cost is a render at $0.10 per started output minute. The same video in a layered tool has a narration track, a music track and maybe a handful of effects, nowhere near 24.
A useful habit is to write the audio plan as a list before you build anything: narration pieces, music, effects. If the list has more than 20 spoken pieces or more than one layered bed, you already know which tool's limit you will meet first. So for typical presenter videos neither cap matters. Check them only if you are building something with dozens of voices, such as a multi-speaker documentary or a game trailer. For those, count your pieces before you pick a tool.
Sources
Related posts
More in Comparisons
- Syren can't use your own avatar or cloned voice; Sume has a handle
Synthesia says own avatars and cloned voices need its regular editor, not Syren. Sume builds an avatar from a photo; voice cloning is app-only.
- Syren's 3-minute cap and 30-second avatar scenes vs Sume's 60 s
Syren runs up to 3 minutes but avatar scenes over 30 s can time out; Sume takes 4-60 s per avatar job. How to plan a 3-minute presenter video.
- WaveSpeed levels come from one top-up; Sume concurrency is plan-only
WaveSpeed sets Bronze to Ultra by a single top-up amount. On Sume a top-up never raises concurrency; the plan does (1, 4, 8, 20). What that means for budgets.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
Written by Sume