Chapter markers from Sume STT sentence segments

Request segmentation mode sentence on Sume STT, get time ranges per sentence, and turn the ones you pick into 0:00-style chapter lines with a 12-line formatter.

5 min readSume
All posts

Sume STT can return sentence segments next to the words. Send segmentation: { mode: "sentence" } with your audio_url to POST /v1/stt-1.0/transcribe, and each segment arrives as { index, text, start, end, duration_seconds }. To build chapter markers, pick the segments that open a topic, and format each start time as m:ss. The segments are time ranges over your original file; no sliced audio is produced.

Request and result

The docs describe segmentation as sentence grouping derived from the returned word timings: words are grouped on terminal punctuation, and unpunctuated runs split on silence. boundary_lead_ms defaults to 70 and carries a little lead past a sentence's last word before the next segment starts, the same default the TTS surface uses. If the provider returns no timed words, the request fails closed with a typed error rather than guessing.

Picking chapter starts

Sentences are not chapters. Two simple approaches work. Either read the segment list and mark the starts by hand, or ask a language model to pick the segments that open a new topic and return the index values. Either way, the timestamps come from the segments, not from a model's memory.

  • Start the list at 0:00 with the opening chapter.
  • Use at least three chapters, each at least ten seconds long.
  • Prefer a segment that starts after a pause over one that starts mid-thought.

Formatting

Given a list of (start_seconds, title) pairs, the lines look like 0:00 Intro. For files over an hour, pad hours as h:mm:ss.

def chapter_lines(chapters):
    lines = []
    for start, title in chapters:
        s = int(start)
        h, m, sec = s // 3600, s % 3600 // 60, s % 60
        stamp = f"{h}:{m:02}:{sec:02}" if h else f"{m}:{sec:02}"
        lines.append(f"{stamp} {title}")
    return "\n".join(lines)

print(chapter_lines([(0, "Intro"), (95.4, "Pricing"), (3725, "Q and A")]))

Limits

The API describes duration_seconds as a reservation hint with a maximum of 10 minutes, so plan on ten-minute jobs for long recordings; shift each segment start by its chunk offset when you merge them; the chunking post shows the merge. Pricing is $0.01 per audio minute. Compare that with streaming recognisers such as MAI-Transcribe-2-Streaming at $0.54 per hour (intro rate, read 2026-10-04), which are built for live speech rather than for chaptering a file. For the poll loop, see jobs and results.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume