Sume TTS segmentation 400: segmentation_requires_word_timestamps

Asking for segmentation.mode sentence without timestamps.words set to true is refused with segmentation_requires_word_timestamps. Turn word timestamps on.

3 min readSume
All posts

A TTS request with segmentation but without timestamps.words: true is refused with a message naming segmentation_requires_word_timestamps. Add "timestamps": { "words": true } beside the segmentation object and send it again.

A valid pair

The two blocks work together: word timestamps are the input to sentence segmentation. segmentation.mode accepts only sentence.

{
  "model": "sonic-3.6",
  "transcript": "First sentence. Second sentence.",
  "voice": { "id": "VOICE_UUID" },
  "timestamps": { "words": true },
  "segmentation": { "mode": "sentence", "emit_audio": false }
}

Segmentation options

boundary_lead_ms is an integer from 0 to 500 and defaults to 70. emit_audio defaults to true and needs a wav or raw output container, so set it to false when you keep the default mp3 and only need the times.

TTS segmentation fields (read 2026-10-03)
FieldAllowedDefault
timestamps.wordstrue or falseNot set
segmentation.modesentenceRequired inside segmentation
segmentation.boundary_lead_ms0 to 50070
segmentation.emit_audiotrue or falsetrue

Why a 400 will not fix itself

A validation refusal does not change on retry. Correct the body before resubmitting, with a new idempotency key if your client treats the failed attempt as final. The boundary lead post explains what the 70 ms lead does to cut points.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume