Sume STT sentence segmentation fails closed when no words are timed

If the speech provider returns no timed words, a Sume STT request with segmentation returns a typed error instead of guessed sentences. Plan for it.

5 min readSume
All posts

When you ask Sume STT for sentence segmentation and the provider returns no timed words, the request does not invent boundaries. It fails closed with a typed error. That is deliberate: sentence segments are derived from word timings, so no timings means no honest segments. Your code should treat that error as a normal branch, for example silent audio or a file with no speech.

Below is what the spec says, and a way to fall back without paying twice.

What does the spec say?

In the Sume OpenAPI spec, segmentation is described as optional sentence segmentation derived from the returned word timings, and fails closed with a typed error if the provider returns no timed words. The only mode is sentence, and boundary_lead_ms accepts 0 to 500, default 70. Word timings are always returned by the job.

Segmentation behaviour, Sume spec read 2026-10-04
CaseWhat happens
Speech with timed wordsSentences are cut from the word timings
Provider returns no timed wordsTyped error, no guessed segments
segmentation omittedTranscript and word timings, no sentence grouping
boundary_lead_ms0 to 500, default 70

When would a file have no timed words?

The usual causes are a recording with no speech, a track that is only music, or a muted channel. Check the source before you blame the API.

  • Play the first few seconds of the file you sent.
  • Confirm the URL points at the audio track you meant.
  • Retry without segmentation if you only need the transcript.

How do you code the fallback?

Submit with segmentation; if the job fails, read the error from the job record and decide. The exact error code string is in the job's error object, so print it rather than hardcoding a guess.

import os, time, requests

H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
B = "https://api.sume.com"

r = requests.post(B + "/v1/stt-1.0/transcribe", headers=H, timeout=60, json={
    "audio_url": os.environ["AUDIO_URL"],
    "segmentation": {"mode": "sentence"},
})
r.raise_for_status()
job_id = r.json()["data"]["job"]["id"]
while True:
    d = requests.get(B + f"/v1/jobs/{job_id}/status", headers=H,
                     timeout=30).json()["data"]
    if d["terminal"]:
        break
    time.sleep(2)
print(d["sume_status"])

What do you build on segments?

Segments feed caption cues and reading-speed checks; see checking caption reading speed from STT segments and an interactive transcript from word timestamps.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume