Text to speech sentence timestamps API: one lip-sync clip each
Sume TTS 1.0 returns gapless sentence segments; with wav output each gets its own audio_url, ready to drive a lip-sync clip. Request shape and the 5 s catch.

To get sentence timestamps from a text-to-speech API on Sume, send timestamps.words: true and segmentation.mode: "sentence" to POST /v1/tts-1.0/generate. The result carries gapless segments[], and when the output container is wav each segment also has its own audio file, so every sentence can drive one lip-sync clip without you cutting audio yourself.
The details come from the TTS 1.0 schema in the OpenAPI document behind the API reference and the lip-sync request schemas in the same document, read 2026-09-29.
What does the request look like?
Send the transcript, a voice selector, a wav container, word timestamps, and sentence segmentation. Segmentation requires timestamps.words: true, and segment audio is only sliced for wav or raw containers.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"transcript": "Welcome back. Today we launch the new line. It ships Friday.",
"avatar_handle": "@your-avatar",
"output_format": { "container": "wav", "sample_rate": 44100, "encoding": "pcm_s16le" },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" }
}'What are the segments guaranteed to be?
The schema promises gapless segments, meaning each segment ends exactly where the next begins, and a default 70 ms lead after the last word of a sentence (boundary_lead_ms, 0 to 500). The next segment absorbs the pause. With wav or raw, emit_audio defaults to true and each segment includes a sample-exact audio_url. With mp3 you get the timings but no segment audio.
| Container | Segment timings | Per-sentence audio_url |
|---|---|---|
wav | Yes | Yes, sample-exact |
raw | Yes | Yes |
mp3 | Yes | No |
Can a segment go straight into a lip-sync request?
Yes, that is what the lip-sync schemas describe: both take a Sume-hosted audio_url, described as typically a TTS segment, at most 10 MB, and non-Sume hosts are rejected. Pair it with one still image or a ready avatar and the output length follows the audio.
Now the catch. POST /v1/minimax/h3-max/lip-sync accepts only 5 to 14.8 seconds of audio and refuses anything outside that window rather than clamping it. A short sentence such as "It ships Friday." will usually be under 5 seconds, so for that route you merge neighbouring sentences into one clip. POST /v1/veed/fabric-1.0 takes a duration_seconds of 1 to 300, so a short sentence fits.
How do I stitch the clips back together?
Because the segments are gapless, the clips rendered from them line up back to back. The Timeline audio page describes a sample-domain join of Sume-hosted audio with no silence at the seams, and its concat takes 1 to 20 ordered parts. Use it for the audio spine and place the lip-sync clips on a timeline in segment order.
When is this not the right shape?
If one long line is what you need, skip the segmenting: a single lip-sync job on the 300-second route handles it. Sentence-per-clip earns its cost when you want to swap a single sentence later or run clips in parallel. For a line that must be cut by time rather than sentence, read lip sync longer than 15 seconds.
Sources
Related posts
More in Developers
- TTS WAV or MP3 for video: keep WAV until the last mux
Sume TTS 1.0 defaults to mp3 at 44100 Hz and 128 kbps. If the audio will be joined, sliced or muxed into video, ask for wav; mp3 adds padding at every edge.
- Let the API pick the video model: sume/auto for vertical UGC clips
Send model sume/auto to POST /v1/videos and Sume picks the family. The response echoes sume/auto and never names the model. When to pin a model instead.
- 400 unknown_parameter on a Format run: read the suggestion
Sume rejects a top-level field it does not know with 400 unknown_parameter and names the likely one. Why webook_url and thread_id fail, and the fix.
- Veo 3.1 personGeneration: allow_adult vs allow_all by mode
Veo 3.1 personGeneration is allow_all for text-to-video, allow_adult for image modes, and allow_adult only in some regions. Sume has no such field.
Written by Sume