MiniMax T2A word subtitles vs Sume TTS timestamps.words

MiniMax T2A offers sentence, word and word_streaming subtitle modes. Sume TTS returns word timestamps through timestamps.words, plus sentence segmentation.

3 min readSume
All posts

MiniMax's T2A HTTP reference lists subtitle options of sentence, word and word_streaming (read 2026-10-03). Sume's TTS endpoints return word timings through timestamps.words, with segmentation.mode: "sentence" for sentence spans, so both can drive karaoke-style captions.

Subtitle options

The MiniMax reference also limits a request to under 10,000 characters, so long scripts need to be split before subtitle files are stitched. Sume allows up to 20,000 characters per request.

Timing options for generated speech (read 2026-10-03)
ItemMiniMax T2ASume TTS
Word timingssubtitle option wordtimestamps.words
Sentence timingssubtitle option sentencesegmentation.mode sentence
Streaming word timingsword_streamingNot documented
Characters per requestUnder 10,00020,000

Sentence boundaries on Sume

With sentence segmentation, boundary_lead_ms defaults to 70, which starts each sentence slightly early so cuts feel tight. emit_audio needs a wav or raw container, because sentence audio is sliced from uncompressed output. The default output is mp3 at 44,100 Hz and 128 kbps.

What to do with the timings

Feed word times to a caption style that highlights the active word. Sume's caption endpoint accepts words and cues directly, so you can skip a second transcription pass when you already generated the speech. The inline captions post lists the cost trade-offs.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume