Sume STT segmentation: boundary_lead_ms 70 default, 0-500 range

Sume STT can return sentence segments with a boundary_lead_ms from 0 to 500, default 70. What it is for when you turn those sentences into caption cues.

4 min readSume
All posts

Sume speech-to-text takes an optional segmentation object with mode sentence and a boundary_lead_ms integer between 0 and 500 that defaults to 70. It returns time ranges for sentences over your own audio, which are a natural starting point for caption cues. Unlike TTS segmentation there is no emit_audio field, because the audio is your file.

The request

The STT body takes audio_url, which must be a public HTTPS URL, plus optional language_code, duration_seconds and segmentation. The schema is strict, so unknown fields are rejected. The segmentation object mirrors the one TTS uses so that a recording you supply and audio Sume synthesises are described the same way.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
FieldType and rangeDefault
audio_urlPublic HTTPS URL, up to 2,048 charactersRequired
language_codeString, 2 to 16 charactersOptional hint
duration_secondsInteger 1 to 600Estimated as 60 if omitted
segmentation.modesentenceOptional object
segmentation.boundary_lead_msInteger 0 to 50070
{
  "audio_url": "https://example.com/talk.wav",
  "language_code": "en",
  "duration_seconds": 240,
  "segmentation": {"mode": "sentence", "boundary_lead_ms": 70}
}

Turning segments into cues

A cue should start a little early so the eye lands on the words as they begin. That is the sensible use of a small lead. Start with the 70 ms default, review a handful of cues against the video, and move toward 150 or 200 only if the text feels late. Beyond a few hundred milliseconds, cues begin to appear before the speaker has started, which reads as a mistake.

What to verify on your own clip

Run the same 60-second clip at 0, 70 and 200 milliseconds and compare the first cue of each sentence against the waveform. Choose the smallest lead that stops cues feeling late. Because the field has a hard ceiling of 500, you can never push a cue half a second early by accident.

Check the response for the exact segment fields before you write parsing code, since this post lists the request side only.

Limits to plan for

Sentences can be long. A 30-word sentence is hard to read in one caption, so split long segments by word count or by pause before you burn them. The segmentation itself does not split by length. Check reading speed too; the captions-versus-script post linked below covers aligning to your own script text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume