Subtitle cues from STT sentence segments and the 70 ms lead

Ask Sume STT for sentence segments and tune boundary_lead_ms (default 70) so each cue holds a little past its last word. Settings and limits.

4 min readSume
All posts

To get subtitle-ready sentences, add segmentation: {"mode": "sentence"} to your Sume STT request. Sume groups the returned words into sentences on terminal punctuation, and splits unpunctuated runs on silence. boundary_lead_ms (0 to 500, default 70) is the milliseconds carried past a sentence's last word before the next segment starts.

Request

Sume STT 1.0 request fields, from the Sume OpenAPI reference, read 2026-10-01.
FieldValueNote
audio_urlpublic HTTPS URLSume media URL preferred
language_codefor example enomit to auto-detect
duration_seconds1 to 600omit and Sume reserves 1 minute
segmentation.modesentencethe only mode
segmentation.boundary_lead_ms0 to 500, default 70same rule as tts-1.0

Turning segments into cues

A cue needs a start, an end and a line. Use each segment's text and times. Keep cues under about 42 characters per line and two lines per cue for readability on phones; that is a house rule, not an API limit. If a sentence is longer, split it at word boundaries using the word timings that always come back with the result.

Failure mode to know

Segmentation fails closed: if the provider returns no timed words, the request returns a typed error rather than made-up cue times. Handle it by retrying without segmentation and building your own cues from the word list. A clip with only music or silence is a plausible cause.

Checklist

  • Set language_code if you know the language, since auto-detect can guess wrong on short clips.
  • Keep clips under 10 minutes per request; split longer audio first.
  • Send an honest duration_seconds.
  • Review the first and last cue by eye; word times are not perfect.
  • If the source is video, detach 16 kHz mono audio first: extracting audio for transcription.

Limits

This returns sentence segments, not styled captions. Burning subtitles into video is a separate step. Price is $0.01 per audio minute; see the API reference. For a click-to-seek view of the same data, see the clickable transcript post.

Related posts

More in Media tools

All Media tools posts

Written by Sume