Text to speech sentence timestamps API: one lip-sync clip each

Sume TTS 1.0 returns gapless sentence segments; with wav output each gets its own audio_url, ready to drive a lip-sync clip. Request shape and the 5 s catch.

4 min readSume
All posts

To get sentence timestamps from a text-to-speech API on Sume, send timestamps.words: true and segmentation.mode: "sentence" to POST /v1/tts-1.0/generate. The result carries gapless segments[], and when the output container is wav each segment also has its own audio file, so every sentence can drive one lip-sync clip without you cutting audio yourself.

The details come from the TTS 1.0 schema in the OpenAPI document behind the API reference and the lip-sync request schemas in the same document, read 2026-09-29.

What does the request look like?

Send the transcript, a voice selector, a wav container, word timestamps, and sentence segmentation. Segmentation requires timestamps.words: true, and segment audio is only sliced for wav or raw containers.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "transcript": "Welcome back. Today we launch the new line. It ships Friday.",
    "avatar_handle": "@your-avatar",
    "output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence" }
  }'

What are the segments guaranteed to be?

The schema promises gapless segments, meaning each segment ends exactly where the next begins, and a default 70 ms lead after the last word of a sentence (boundary_lead_ms, 0 to 500). The next segment absorbs the pause. With wav or raw, emit_audio defaults to true and each segment includes a sample-exact audio_url. With mp3 you get the timings but no segment audio.

What sentence segmentation returns by container, read 2026-09-29.
ContainerSegment timingsPer-sentence audio_url
wavYesYes, sample-exact
rawYesYes
mp3YesNo

Can a segment go straight into a lip-sync request?

Yes, that is what the lip-sync schemas describe: both take a Sume-hosted audio_url, described as typically a TTS segment, at most 10 MB, and non-Sume hosts are rejected. Pair it with one still image or a ready avatar and the output length follows the audio.

Now the catch. POST /v1/minimax/h3-max/lip-sync accepts only 5 to 14.8 seconds of audio and refuses anything outside that window rather than clamping it. A short sentence such as "It ships Friday." will usually be under 5 seconds, so for that route you merge neighbouring sentences into one clip. POST /v1/veed/fabric-1.0 takes a duration_seconds of 1 to 300, so a short sentence fits.

How do I stitch the clips back together?

Because the segments are gapless, the clips rendered from them line up back to back. The Timeline audio page describes a sample-domain join of Sume-hosted audio with no silence at the seams, and its concat takes 1 to 20 ordered parts. Use it for the audio spine and place the lip-sync clips on a timeline in segment order.

When is this not the right shape?

If one long line is what you need, skip the segmenting: a single lip-sync job on the 300-second route handles it. Sentence-per-clip earns its cost when you want to swap a single sentence later or run clips in parallel. For a line that must be cut by time rather than sentence, read lip sync longer than 15 seconds.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume