Lip-sync audio under 5 seconds: join two TTS lines with audio concat
MiniMax H3 Max lip-sync on Sume wants 5 to 14.8 seconds of audio. Join two short TTS lines with Timeline audio concat for $0.01 instead of padding.

If your TTS line is under 5 seconds, it is too short for the H3 Max lip-sync route on Sume, which takes Sume-hosted audio of 5 to 14.8 seconds. Join it to the next line with POST /v1/timeline-1.0/audio using operation: "concat". The join is sample-exact with no re-synthesis and costs $0.01 flat per job.
The window is in the lip-sync route description of the Sume API reference, and the concat rules are in the Timeline audio docs (both read 2026-10-06).
How do I join the lines?
Generate each line with TTS as WAV (output_format wav, pcm_s16le) so the join stays sample-exact; the docs say to keep WAV if the file drives lip-sync. Then send parts[] with up to 20 files, each { url } with an optional source_in and duration. All the URLs have to be audio on this workspace's media.sume.com and have the same channel layout, and an Idempotency-Key is required.
- The result is
kind: timeline_audiowith oneaudio_url,duration_secondsandsegments[]. - Check that
duration_secondsis at least 5 and at most 14.8. - Pass the new
audio_urland aduration_secondsin 5 to 14.8 to the lip-sync route.
What does it cost?
The join is $0.01. H3 Max is billed per second of audio. Check the live rate for your resolution in GET /v1/catalog before you queue a batch. A line that is a little too short is cheaper to pad with the next sentence than to send to a model that refuses it.
What if it is over 14.8 seconds?
The API refuses a duration_seconds outside 5 to 14.8 and never clamps it. Split the audio at a sentence end with the same endpoint and operation: "split", and send each part to the lip-sync route separately.
Sources
More in Sume Avatar 1.0
- One avatar handle, three ratios: the same face in 9:16, 1:1, 16:9
Keep one AI presenter across Shorts, feed and webinar slots: reuse a single avatar_handle and change only aspect_ratio on each Sume talking-video call.
- Pipecat avatar TTFB metrics vs timing a Sume avatar job
Pipecat's Tavus service reports TTFB from TTSStartedFrame and BotStartedSpeakingFrame. A Sume avatar job has no first byte to time: measure submit to completed.
- Pipecat HeyGen LiveAvatar: duration by tier vs a 60 second Sume job
In Pipecat, HeyGen LiveAvatar session length depends on your subscription tier. On Sume the length rule is fixed: 4 to 60 seconds per job, priced per second.
- Pipecat Simli max_session_length vs Sume's 4 to 60 second clip
Simli in Pipecat caps a live session with max_session_length and max_idle_time. Sume has no session: an avatar job is a 4 to 60 second file.
Written by Sume