Syren mixes the first 24 audio tracks; Sume's timeline takes 20 parts

Syren mixes only the first 24 audio tracks in an export. Sume's timeline takes one spine in up to 20 gapless parts plus one bed. What fits where.

4 min readSume
All posts

Syren's help article says it mixes the first 24 audio tracks into your exported video; anything past 24 is not mixed (read 2026-10-11). Sume's Timeline 1.0 has a different model: one audio spine, which can be assembled from up to 20 gapless parts, plus one optional soundtrack bed underneath.

The Syren fact comes from How do I create a video with Syren?, read 2026-10-11. The Sume facts come from the Timeline 1.0 and Timeline audio pages. This is a limits comparison, not a quality ranking.

Two different audio models

Syren is a layered mixer: narration, music and effects are separate tracks that get summed, and Synthesia documents a cap on how many are included. Sume's timeline is a sequence: one spine decides the length of the output and the video slots are laid against it by start time. Layering exists only in one place, the soundtrack bed, which can loop, fade out for up to 10 seconds and duck under speech by 0 to 20 dB when the spine has real audio.

What each limit means in practice

For a talking presenter with a bed of music, both systems fit easily. The limit only bites when you assemble many voices or effects.

Two more timeline details help when you plan. The audio length you declare is the output length, and video slots may stop at most half a second before the end of the spine. Stills are static holds in the timeline, so a slide-style cutaway is a still with a duration rather than an animation.

Audio limits as documented (Synthesia help page read 2026-10-11; Sume Timeline 1.0 docs)
QuestionSyrenSume Timeline 1.0
Tracks in the final mixFirst 24 audio tracksOne spine plus one soundtrack bed
Pieces in the spineNot statedUp to 20 gapless parts
Output lengthUp to 3 minutes1 to 1,800 seconds of audio
Music under speechA track among the 24soundtrack with duck_db 0-20 and fade_out up to 10 s
Silent stretchNot statedaudio.mode silence for a declared length

When you would hit Sume's limit

A 20-part spine covers a 20-segment voice-over. If you have more than 20 spoken pieces, merge some of them first with the separate timeline audio job, which produces a reusable merged file at a fixed price per job (the catalog lists it at $0.01), then use that file as the spine. Because the join is sample-domain with no re-TTS, the voice does not change at the seams.

Sound effects are the real gap. Sume's timeline does not take a stack of effect tracks, so if your video needs many layered sounds, build the effect bed outside the timeline and feed it in as the single soundtrack. If the layered mix is the point of your project, Syren's 24-track model is closer to what you want.

A talking-head example

Take a 90-second founder update. On Sume it is two avatar jobs of 45 seconds each, both spoken by the avatar, joined by detaching each clip's audio into parts and rendering one timeline with an optional music bed ducked under the voice. That is two parts out of 20 and one bed, nowhere near a limit, and the extra timeline cost is a render at $0.10 per started output minute. The same video in a layered tool has a narration track, a music track and maybe a handful of effects, nowhere near 24.

A useful habit is to write the audio plan as a list before you build anything: narration pieces, music, effects. If the list has more than 20 spoken pieces or more than one layered bed, you already know which tool's limit you will meet first. So for typical presenter videos neither cap matters. Check them only if you are building something with dozens of voices, such as a multi-speaker documentary or a game trailer. For those, count your pieces before you pick a tool.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume