Sume STT sentence segmentation fails closed when no words are timed
If the speech provider returns no timed words, a Sume STT request with segmentation returns a typed error instead of guessed sentences. Plan for it.

When you ask Sume STT for sentence segmentation and the provider returns no timed words, the request does not invent boundaries. It fails closed with a typed error. That is deliberate: sentence segments are derived from word timings, so no timings means no honest segments. Your code should treat that error as a normal branch, for example silent audio or a file with no speech.
Below is what the spec says, and a way to fall back without paying twice.
What does the spec say?
In the Sume OpenAPI spec, segmentation is described as optional sentence segmentation derived from the returned word timings, and fails closed with a typed error if the provider returns no timed words. The only mode is sentence, and boundary_lead_ms accepts 0 to 500, default 70. Word timings are always returned by the job.
| Case | What happens |
|---|---|
| Speech with timed words | Sentences are cut from the word timings |
| Provider returns no timed words | Typed error, no guessed segments |
| segmentation omitted | Transcript and word timings, no sentence grouping |
| boundary_lead_ms | 0 to 500, default 70 |
When would a file have no timed words?
The usual causes are a recording with no speech, a track that is only music, or a muted channel. Check the source before you blame the API.
- Play the first few seconds of the file you sent.
- Confirm the URL points at the audio track you meant.
- Retry without
segmentationif you only need the transcript.
How do you code the fallback?
Submit with segmentation; if the job fails, read the error from the job record and decide. The exact error code string is in the job's error object, so print it rather than hardcoding a guess.
import os, time, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"
r = requests.post(B + "/v1/stt-1.0/transcribe", headers=H, timeout=60, json={
"audio_url": os.environ["AUDIO_URL"],
"segmentation": {"mode": "sentence"},
})
r.raise_for_status()
job_id = r.json()["data"]["job"]["id"]
while True:
d = requests.get(B + f"/v1/jobs/{job_id}/status", headers=H,
timeout=30).json()["data"]
if d["terminal"]:
break
time.sleep(2)
print(d["sume_status"])
What do you build on segments?
Segments feed caption cues and reading-speed checks; see checking caption reading speed from STT segments and an interactive transcript from word timestamps.
Sources
Related posts
More in Developers
- Test a Sume STT webhook locally: webhook_url must be public HTTPS
Sume rejects localhost, private-network and non-HTTPS webhook_url values. Put a tunnel in front of your dev server, or poll while you build.
- Authenticate the Sume CLI on a CI runner without a browser login
On CI, skip sume login: install the CLI, run sume auth setup with an API key from a secret, and confirm with sume auth status before any job step.
- sume/auto for a former Sora feature: when to pin a model
Sume's sume/auto picks a family and never says which. Good for general clips, wrong when a brand needs one look. How to choose between auto and a pinned id.
- sume/auto for images: no model named, no seed, so pin ids for brand
sume/auto picks an image family and never says which, and there is no seed. When Auto is fine, and when to pin an id like GPT Image 2.5.
Written by Sume