STT with no punctuation: Sume still cuts sentence segments on silence
Sume STT's sentence segmentation splits on terminal punctuation, and on silence when a run has none. What boundary_lead_ms 70 does and a Python call.

Some speech never gets punctuation: a rambling voice memo, a meeting with crosstalk, a language where the recognizer returns few full stops. Sume's STT API can still hand back sentence-sized pieces. Per the schema, segmentation: {mode: "sentence"} groups words into sentences on terminal punctuation, and splits unpunctuated runs on silence.
The pieces you get
Word timings are always returned, so there is no flag for them. Add segmentation and you also receive sentence segments derived from those word times. boundary_lead_ms (0 to 500, default 70) is how many milliseconds past a sentence's last word the segment runs before the next begins, the same rule sume/tts-1.0 uses.
Fails closed
If the provider returns no timed words, the segmentation request does not guess. Sume returns a typed error. That keeps you from building captions on invented timings.
A request
The audio must be a public HTTPS URL, up to 600 seconds. Sending duration_seconds lets Sume reserve the right amount; without it, Sume reserves one minute. Price is $0.01 per audio minute.
import os
import requests
resp = requests.post(
"https://api.sume.com/v1/stt-1.0/transcribe",
headers={
"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"Idempotency-Key": "memo-001",
},
json={
"audio_url": "https://media.sume.com/example/memo.wav",
"duration_seconds": 95,
"segmentation": {"mode": "sentence", "boundary_lead_ms": 70},
},
timeout=30,
)
print(resp.status_code, resp.json()["data"]["job"]["id"])
Tuning
Raise boundary_lead_ms toward 200 if cuts clip the ends of words; lower it to 0 for tight cuts. A 95-second memo costs about 2 cents at $0.01 per audio minute. Check a few segments by ear. Silence-based splits follow pauses, not grammar, so a long breathless sentence stays one segment.
Related posts
More in Developers
- STT words into caption words: rename word to text, stay under 60 s
Sume STT returns words as {word, start, end}; the caption job wants {text, start, end}, end above start, 60 s or less, 1200 words at most. Map and filter.
- Submit Sume jobs from Go with a bounded pool and idempotency keys
Bound a Go worker pool to your workspace in-flight headroom, send a stable Idempotency-Key per item, and stop treating a full queue as a crash. Stdlib only.
- Sume 400 unsupported_capability on 4:5: fall back to 3:4 in Python
A 4:5 request on a Sume video model returns 400 unsupported_capability before billing. Catch it in Python, resubmit at 3:4, then crop to Meta's 4:5 Feed ratio.
- Sume 429: read error.details.scope, and rate_limit_unavailable
A Sume 429 names the budget in error.details.scope and gives retry_after_seconds. A degraded rate_limit_unavailable 429 is a different case. Python handler.
Written by Sume