Subtitle cues from STT sentence segments and the 70 ms lead
Ask Sume STT for sentence segments and tune boundary_lead_ms (default 70) so each cue holds a little past its last word. Settings and limits.

To get subtitle-ready sentences, add segmentation: {"mode": "sentence"} to your Sume STT request. Sume groups the returned words into sentences on terminal punctuation, and splits unpunctuated runs on silence. boundary_lead_ms (0 to 500, default 70) is the milliseconds carried past a sentence's last word before the next segment starts.
Request
| Field | Value | Note |
|---|---|---|
| audio_url | public HTTPS URL | Sume media URL preferred |
| language_code | for example en | omit to auto-detect |
| duration_seconds | 1 to 600 | omit and Sume reserves 1 minute |
| segmentation.mode | sentence | the only mode |
| segmentation.boundary_lead_ms | 0 to 500, default 70 | same rule as tts-1.0 |
Turning segments into cues
A cue needs a start, an end and a line. Use each segment's text and times. Keep cues under about 42 characters per line and two lines per cue for readability on phones; that is a house rule, not an API limit. If a sentence is longer, split it at word boundaries using the word timings that always come back with the result.
Failure mode to know
Segmentation fails closed: if the provider returns no timed words, the request returns a typed error rather than made-up cue times. Handle it by retrying without segmentation and building your own cues from the word list. A clip with only music or silence is a plausible cause.
Checklist
- Set
language_codeif you know the language, since auto-detect can guess wrong on short clips. - Keep clips under 10 minutes per request; split longer audio first.
- Send an honest
duration_seconds. - Review the first and last cue by eye; word times are not perfect.
- If the source is video, detach 16 kHz mono audio first: extracting audio for transcription.
Limits
This returns sentence segments, not styled captions. Burning subtitles into video is a separate step. Price is $0.01 per audio minute; see the API reference. For a click-to-seek view of the same data, see the clickable transcript post.
Related posts
More in Media tools
- Threads API video requirements: codecs, 5 minutes, 1 GB, 23-60 fps
Threads API video limits as of October 2026: MP4 or MOV, H264 or HEVC, 23-60 fps, 5 minutes, 1 GB. A preflight that checks each one against a Sume probe.
- TikTok trending video transcripts via API: what summary_mode returns
Can you get a transcript with Sume's trending video search? summary_mode transcript returns metadata plus an unsupported warning. What to do for the words.
- Trim an AI clip first, or use source_in on the timeline?
For a shot that is too long, set source_in and duration on the Timeline 1.0 slot. Use video_trim ($0.02 per job) only when another step needs a new MP4 file.
- Put a product shot above a UGC creator clip: compose stack
Timeline compose stacks one Sume-hosted still over one video into a single MP4. Set ratio for the still's share. The public rate is $0.02 per job.
Written by Sume