STT sentence segmentation: unpunctuated speech split on silence

Sume STT 1.0 segmentation mode sentence returns gapless sentence segments, splitting unpunctuated runs on silence. Options, boundary_lead_ms and caveats.

5 min readSume
All posts

What you get

Sume STT 1.0 always returns word timings. If you also send segmentation with mode: "sentence", the response carries sentence segments[]. The API reference says they are gapless and ordered, so each segment's end equals the next segment's start, and times range over the submitted audio_url: no sliced audio files are produced.

The same boundary rule as TTS applies, with a default boundary_lead_ms of 70.

When there is no punctuation

Speech recognition does not always return punctuation, especially for casual or fast talk. The docs say unpunctuated runs are split on silence, so you still get segments you can use as caption lines.

Segmentation options (read 2026-10-03)
FieldValueEffect
segmentation.modesentenceReturn segments[]
segmentation.boundary_lead_ms0 to 500, default 70Shift where a boundary falls after a word
duration_seconds1 to 600Reservation hint; absent reserves 1 minute
language_codeoptionalOmit for auto-detect

A request

Send it to POST /v1/stt-1.0/transcribe. The price is $0.01 per audio minute, and a call takes up to 10 minutes of audio.

{
  "audio_url": "https://example.com/talk.mp3",
  "duration_seconds": 240,
  "segmentation": { "mode": "sentence" }
}

What to do with the segments

  • Map each one to a caption cue or a subtitle line.
  • Cut talk blocks yourself; the docs say no sliced files are produced, so use Timeline audio to slice.
  • Check long segments; a monologue without pauses can yield a long one, so cap length in your own code.
  • Remember speaker labels are not exposed, because diarization is fixed server-side.

Using boundary_lead_ms

boundary_lead_ms ranges from 0 to 500 with a default of 70, the same rule and default as TTS. It controls how far past the last word a boundary is placed, which matters when you cut audio: too small a lead can clip a word's tail, too large can include the next breath.

Start with the default and change it only when you hear clipped endings.

Takeaway

Ask for sentence mode when you need lines rather than words. It works without punctuation, but check the longest segments before they become captions.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume