Chapter markers from Sume STT sentence segments
Request segmentation mode sentence on Sume STT, get time ranges per sentence, and turn the ones you pick into 0:00-style chapter lines with a 12-line formatter.

Sume STT can return sentence segments next to the words. Send segmentation: { mode: "sentence" } with your audio_url to POST /v1/stt-1.0/transcribe, and each segment arrives as { index, text, start, end, duration_seconds }. To build chapter markers, pick the segments that open a topic, and format each start time as m:ss. The segments are time ranges over your original file; no sliced audio is produced.
Request and result
The docs describe segmentation as sentence grouping derived from the returned word timings: words are grouped on terminal punctuation, and unpunctuated runs split on silence. boundary_lead_ms defaults to 70 and carries a little lead past a sentence's last word before the next segment starts, the same default the TTS surface uses. If the provider returns no timed words, the request fails closed with a typed error rather than guessing.
Picking chapter starts
Sentences are not chapters. Two simple approaches work. Either read the segment list and mark the starts by hand, or ask a language model to pick the segments that open a new topic and return the index values. Either way, the timestamps come from the segments, not from a model's memory.
- Start the list at 0:00 with the opening chapter.
- Use at least three chapters, each at least ten seconds long.
- Prefer a segment that starts after a pause over one that starts mid-thought.
Formatting
Given a list of (start_seconds, title) pairs, the lines look like 0:00 Intro. For files over an hour, pad hours as h:mm:ss.
def chapter_lines(chapters):
lines = []
for start, title in chapters:
s = int(start)
h, m, sec = s // 3600, s % 3600 // 60, s % 60
stamp = f"{h}:{m:02}:{sec:02}" if h else f"{m}:{sec:02}"
lines.append(f"{stamp} {title}")
return "\n".join(lines)
print(chapter_lines([(0, "Intro"), (95.4, "Pricing"), (3725, "Q and A")]))Limits
The API describes duration_seconds as a reservation hint with a maximum of 10 minutes, so plan on ten-minute jobs for long recordings; shift each segment start by its chunk offset when you merge them; the chunking post shows the merge. Pricing is $0.01 per audio minute. Compare that with streaming recognisers such as MAI-Transcribe-2-Streaming at $0.54 per hour (intro rate, read 2026-10-04), which are built for live speech rather than for chaptering a file. For the poll loop, see jobs and results.
Sources
Related posts
More in Developers
- Check Sume artifact size, width and duration before you download
Read size_bytes, width, height and duration_ms from the job result and reject an unexpected artifact before spending bandwidth. Node 18 TypeScript sample.
- Choose an image model in code from the Sume catalog's parameters
Filter GET /v1/images/models by what a request needs (references, ratio, transparency), then rank the matches by endpoint price. Python script for Sume.
- Claude batch custom_id is 64 characters: keep SKU keys valid
Anthropic batch custom_id allows 1 to 64 letters, digits, underscore and hyphen. How to sanitize SKUs, avoid collisions and carry the key into Sume input.
- Claude Code 2.1.285 lists WebSocket MCP servers; Sume uses HTTP
Claude Code 2.1.285 shows WebSocket MCP servers in claude mcp list. Sume's hosted MCP is a remote HTTP server, added with --transport http.
Written by Sume