Sync sync_mode bounce, cut_off, silence, remap vs Sume lengths
Sync offers bounce, cut_off, silence and remap when audio and video lengths differ. Sume has no such option: trim or split audio to the accepted window.

Sync's changelog adds a sync mode selector for when audio and video have different durations: bounce loops the shorter media back and forth, cut_off trims the excess, silence pads the gap and remap stretches to fit. Sume's talking-video calls have no equivalent option, so you make the audio fit the accepted window before you submit.
Sync facts are from its changelog; Sume limits from the OpenAPI document and the Models and Avatar videos docs, read 2026-10-01.
What does each Sync mode do?
| Mode | Behavior when lengths differ |
|---|---|
| bounce | Loops the shorter media back and forth |
| cut_off | Trims the excess |
| silence | Pads the gap |
| remap | Stretches to fit |
What are Sume's length limits?
The Fabric route takes a still and audio. The audio must be on the Sume media host and at most 10 MB, and duration_seconds (used to reserve credits) has a maximum of 300. MiniMax H3 Max Lip Sync uses the same still plus audio body with audio of 5 to 14.8 seconds. Avatar videos accept a script or scene plan Sume estimates at 4-60 seconds inclusive; shorter or longer scripts must be shortened or split into multiple jobs.
So how do I handle a mismatch?
The docs describe no remap, bounce or pad option, so you do the equivalent yourself. To trim, cut the audio before upload. To pad, add silence to the audio file. To fit a long script, split it into parts under the limit and make one job per part, as in lip sync clips longer than 15 seconds. Send an accurate measured duration_seconds, since it drives the credit reservation.
Should I stretch audio to fit?
That is your call, and Sume does not do it for you. If you want a slower or faster read, set the TTS speed multiplier (range 0.6 to 1.5) when you generate the speech instead of time-stretching a finished file. Check the result by ear.
Sources
Related posts
More in Developers
- Synthesia Billing API credit usage vs Sume's /v1/usage ledger
Synthesia added a Billing API endpoint for credit usage. Sume's equivalent is GET /v1/usage: reservations, captures and refunds, filterable by run or job.
- Synthesia dubbing file size limit: 5 GB or 2.5 hours
Synthesia dubbing accepts uploads up to 5 GB or 2.5 hours, 4K, in .mp4, .webm or .mov. Sume's limits sit at the audio stage: a 10 MB audio_url and 300 s.
- Synthesia Interactive Avatar and LiveKit: Sume's job-based pieces
Synthesia's Interactive Avatar API runs on a LiveKit plugin with your own LLM and STT. Sume offers job pieces: TTS, speech to text and still-plus-audio clips.
- Synthesia rejected status: moderation, error, and your handler
Synthesia marks a moderated video rejected, separate from error. Sume reports content_policy_rejected as a public_reason; an unfunded run is a 402 instead.
Written by Sume