Chapter timestamps for narrated audio from concat segment offsets
Join one TTS file per chapter with timeline audio concat, then turn the returned segments[] start offsets into mm:ss chapter lines with a short Python script.

Why this works
If you narrate an audiobook, lecture series or long explainer one chapter per text-to-speech job, the concat job that joins them also gives you the chapter timestamps. Its result lists segments[] with index, start and duration_seconds for every part, so a chapter list is a formatting step, not a transcription step.
The join is sample-domain with no silence at the seams, according to the timeline audio docs, which is why the offsets are exact: part 3 starts where parts 1 and 2 end.
The inputs and outputs
Generate one TTS job per chapter. Pass the finished files, in order, as parts[] on POST /v1/timeline-1.0/audio with operation: "concat". Up to 20 parts per join. Read the result from the job envelope at GET /v1/jobs/:id/result. The kind is timeline_audio, and it holds audio_url, duration_seconds and segments[].
The script below takes a result-shaped dictionary and prints chapter lines. The numbers in the sample are made up to show the format; replace them with your own result.
chapters = ["Introduction", "Setup", "First build", "Shipping"]
result = {
"segments": [
{"index": 0, "start": 0.0, "duration_seconds": 74.2},
{"index": 1, "start": 74.2, "duration_seconds": 191.5},
{"index": 2, "start": 265.7, "duration_seconds": 402.8},
{"index": 3, "start": 668.5, "duration_seconds": 120.0},
]
}
def stamp(seconds):
s = int(seconds)
h, rem = divmod(s, 3600)
m, sec = divmod(rem, 60)
return f"{h}:{m:02d}:{sec:02d}" if h else f"{m}:{sec:02d}"
for seg in sorted(result["segments"], key=lambda x: x["index"]):
print(stamp(seg["start"]), chapters[seg["index"]])What the script prints
Run it and the first line is 0:00 Introduction, then 1:14 Setup, and so on. Whole seconds are truncated, which is the safe direction for a chapter marker: the marker never lands after the audio it names.
Two things to get right
Chapters from segments depend on one rule: every chapter must be its own part. If you join a chapter that you cut up into several TTS jobs, join those first, then join the chapters, or the offsets you read will be per piece, not per chapter. With more than 20 chapters, split into two joins and add the first join's duration_seconds to every offset of the second.
Cost of the whole book
Each concat is a flat $0.01 job, per the docs, and there is no provider inference. The text-to-speech jobs are the real cost. A 3,000-character chapter is 3,000 characters at $47.50 per million, 14.25 cents, billed as 15 cents. Eight such chapters cost $1.20 in TTS plus $0.01 for the join. The 1,200-second cap on a single TTS job is the other reason to work chapter by chapter; the audiobook chapters post covers how to pick the break points.
Format and cap checks
- Keep WAV for all parts if you will join the result again; MP3 adds priming padding at each edge on every re-encode, according to the docs.
- Parts need the same channel layout, or the job fails with
audio_parts_channel_mismatch. - The joined output is capped at 1800 seconds, so a three-hour book is several joined files, not one.
Turning offsets into a chapter list
Each entry in segments[] has an index, a start and a duration_seconds, so the chapter list is the starts in order. Format each start as minutes and seconds, label it with the title of the part at the same index, and the first chapter begins at 0.
Keep your own list of titles in the same order as the parts[] you sent, since the segments are matched by position. If you reorder parts, reorder the titles before you build the list.
A platform that reads chapters from a description usually expects the first stamp at 0:00 and each one on its own line. Check the format of the platform you publish to before you rely on a layout.
- Chapter start = segment
start. - Titles follow the order of
parts[]. - First stamp at 0:00.
Sources
Related posts
More in Developers
- Check an Omni edit kept the rest of the clip: video-frames pairs
Compare stills from the source and the edited clip at the same timestamps with video-frames on Sume. A script that submits both extracts, plus what to look for.
- Check an Omni edit's length with video-inspect before a timeline join
An edit should follow the source length. Confirm it with a probe-only video-inspect call before the clip goes into a Timeline render, with a Python read.
- Check duration, resolution, ratio against /v1/videos/models in Node
A Sora-era request will not fit every Sume model. A Node script reads GET /v1/videos/models and lists what the model rejects before you pay for a job.
- communication.webhook_url 400: HTTPS, public host, 2048 chars
A communication.webhook_url that is not public HTTPS, is over 2048 characters, or points at localhost or a private network returns 400 invalid_request.
Written by Sume