Sume STT segmentation: boundary_lead_ms 70 default, 0-500 range
Sume STT can return sentence segments with a boundary_lead_ms from 0 to 500, default 70. What it is for when you turn those sentences into caption cues.

Sume speech-to-text takes an optional segmentation object with mode sentence and a boundary_lead_ms integer between 0 and 500 that defaults to 70. It returns time ranges for sentences over your own audio, which are a natural starting point for caption cues. Unlike TTS segmentation there is no emit_audio field, because the audio is your file.
The request
The STT body takes audio_url, which must be a public HTTPS URL, plus optional language_code, duration_seconds and segmentation. The schema is strict, so unknown fields are rejected. The segmentation object mirrors the one TTS uses so that a recording you supply and audio Sume synthesises are described the same way.
| Field | Type and range | Default |
|---|---|---|
| audio_url | Public HTTPS URL, up to 2,048 characters | Required |
| language_code | String, 2 to 16 characters | Optional hint |
| duration_seconds | Integer 1 to 600 | Estimated as 60 if omitted |
| segmentation.mode | sentence | Optional object |
| segmentation.boundary_lead_ms | Integer 0 to 500 | 70 |
{
"audio_url": "https://example.com/talk.wav",
"language_code": "en",
"duration_seconds": 240,
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70}
}Turning segments into cues
A cue should start a little early so the eye lands on the words as they begin. That is the sensible use of a small lead. Start with the 70 ms default, review a handful of cues against the video, and move toward 150 or 200 only if the text feels late. Beyond a few hundred milliseconds, cues begin to appear before the speaker has started, which reads as a mistake.
What to verify on your own clip
Run the same 60-second clip at 0, 70 and 200 milliseconds and compare the first cue of each sentence against the waveform. Choose the smallest lead that stops cues feeling late. Because the field has a hard ceiling of 500, you can never push a cue half a second early by accident.
Check the response for the exact segment fields before you write parsing code, since this post lists the request side only.
Limits to plan for
Sentences can be long. A 30-word sentence is hard to read in one caption, so split long segments by word count or by pause before you burn them. The segmentation itself does not split by length. Check reading speed too; the captions-versus-script post linked below covers aligning to your own script text.
Sources
Related posts
More in Developers
- Sume 401 halfway through a batch: stop every worker, do not retry
A 401 on a Sume submit means a missing, malformed or revoked key. The SDK does not retry it, and neither should you. Python sample that keeps job ids.
- A 5xx on a Sume paid submit never proves no job: retry with the key
On Sume, only a validation, authorization or balance error proves a paid create was refused. A 5xx does not, so retry with the same Idempotency-Key.
- sume/auto 400 unsupported_capability: 11 s, 2 s, 480p and silent audio
sume/auto fails closed on 2 s, 11 s, 480p, 768p and generate_audio false. The request is rejected before any provider call. Here is what to send instead.
- AI video client timeout: the Sume job still runs and still bills
A timeout on your HTTP client does not cancel a Sume job. Save the job id, poll the polling_url, and cancel only before generation starts. Python example.
Written by Sume