Sume TTS segmentation 400: segmentation_requires_word_timestamps
Asking for segmentation.mode sentence without timestamps.words set to true is refused with segmentation_requires_word_timestamps. Turn word timestamps on.

A TTS request with segmentation but without timestamps.words: true is refused with a message naming segmentation_requires_word_timestamps. Add "timestamps": { "words": true } beside the segmentation object and send it again.
A valid pair
The two blocks work together: word timestamps are the input to sentence segmentation. segmentation.mode accepts only sentence.
{
"model": "sonic-3.6",
"transcript": "First sentence. Second sentence.",
"voice": { "id": "VOICE_UUID" },
"timestamps": { "words": true },
"segmentation": { "mode": "sentence", "emit_audio": false }
}Segmentation options
boundary_lead_ms is an integer from 0 to 500 and defaults to 70. emit_audio defaults to true and needs a wav or raw output container, so set it to false when you keep the default mp3 and only need the times.
| Field | Allowed | Default |
|---|---|---|
| timestamps.words | true or false | Not set |
| segmentation.mode | sentence | Required inside segmentation |
| segmentation.boundary_lead_ms | 0 to 500 | 70 |
| segmentation.emit_audio | true or false | true |
Why a 400 will not fix itself
A validation refusal does not change on retry. Correct the body before resubmitting, with a new idempotency key if your client treats the failed attempt as final. The boundary lead post explains what the 70 ms lead does to cut points.
Sources
Related posts
More in Media tools
- Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
- Upload 15 Shorts at once in Studio: probe every file first
Studio takes up to 15 Shorts per upload, each up to 3 minutes and square or vertical. Check duration and frame shape with video-inspect before you queue.
- Upscale an old Sora download: the limits of Sume's video upscaler
Sora files you saved can be upscaled. Sume's video-upscale model takes a scale from 1.1 to 4, up to 30 seconds, and a fast, standard or pro tier.
Written by Sume