Sync sync_mode bounce, cut_off, silence, remap vs Sume lengths

Sync offers bounce, cut_off, silence and remap when audio and video lengths differ. Sume has no such option: trim or split audio to the accepted window.

4 min readSume
All posts

Sync's changelog adds a sync mode selector for when audio and video have different durations: bounce loops the shorter media back and forth, cut_off trims the excess, silence pads the gap and remap stretches to fit. Sume's talking-video calls have no equivalent option, so you make the audio fit the accepted window before you submit.

Sync facts are from its changelog; Sume limits from the OpenAPI document and the Models and Avatar videos docs, read 2026-10-01.

What does each Sync mode do?

Sync sync modes from its changelog, read 2026-10-01.
ModeBehavior when lengths differ
bounceLoops the shorter media back and forth
cut_offTrims the excess
silencePads the gap
remapStretches to fit

What are Sume's length limits?

The Fabric route takes a still and audio. The audio must be on the Sume media host and at most 10 MB, and duration_seconds (used to reserve credits) has a maximum of 300. MiniMax H3 Max Lip Sync uses the same still plus audio body with audio of 5 to 14.8 seconds. Avatar videos accept a script or scene plan Sume estimates at 4-60 seconds inclusive; shorter or longer scripts must be shortened or split into multiple jobs.

So how do I handle a mismatch?

The docs describe no remap, bounce or pad option, so you do the equivalent yourself. To trim, cut the audio before upload. To pad, add silence to the audio file. To fit a long script, split it into parts under the limit and make one job per part, as in lip sync clips longer than 15 seconds. Send an accurate measured duration_seconds, since it drives the credit reservation.

Should I stretch audio to fit?

That is your call, and Sume does not do it for you. If you want a slower or faster read, set the TTS speed multiplier (range 0.6 to 1.5) when you generate the speech instead of time-stretching a finished file. Check the result by ear.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume