Chapter timestamps for narrated audio from concat segment offsets

Join one TTS file per chapter with timeline audio concat, then turn the returned segments[] start offsets into mm:ss chapter lines with a short Python script.

5 min readSume
All posts

Why this works

If you narrate an audiobook, lecture series or long explainer one chapter per text-to-speech job, the concat job that joins them also gives you the chapter timestamps. Its result lists segments[] with index, start and duration_seconds for every part, so a chapter list is a formatting step, not a transcription step.

The join is sample-domain with no silence at the seams, according to the timeline audio docs, which is why the offsets are exact: part 3 starts where parts 1 and 2 end.

The inputs and outputs

Generate one TTS job per chapter. Pass the finished files, in order, as parts[] on POST /v1/timeline-1.0/audio with operation: "concat". Up to 20 parts per join. Read the result from the job envelope at GET /v1/jobs/:id/result. The kind is timeline_audio, and it holds audio_url, duration_seconds and segments[].

The script below takes a result-shaped dictionary and prints chapter lines. The numbers in the sample are made up to show the format; replace them with your own result.

chapters = ["Introduction", "Setup", "First build", "Shipping"]

result = {
    "segments": [
        {"index": 0, "start": 0.0, "duration_seconds": 74.2},
        {"index": 1, "start": 74.2, "duration_seconds": 191.5},
        {"index": 2, "start": 265.7, "duration_seconds": 402.8},
        {"index": 3, "start": 668.5, "duration_seconds": 120.0},
    ]
}

def stamp(seconds):
    s = int(seconds)
    h, rem = divmod(s, 3600)
    m, sec = divmod(rem, 60)
    return f"{h}:{m:02d}:{sec:02d}" if h else f"{m}:{sec:02d}"

for seg in sorted(result["segments"], key=lambda x: x["index"]):
    print(stamp(seg["start"]), chapters[seg["index"]])

What the script prints

Run it and the first line is 0:00 Introduction, then 1:14 Setup, and so on. Whole seconds are truncated, which is the safe direction for a chapter marker: the marker never lands after the audio it names.

Two things to get right

Chapters from segments depend on one rule: every chapter must be its own part. If you join a chapter that you cut up into several TTS jobs, join those first, then join the chapters, or the offsets you read will be per piece, not per chapter. With more than 20 chapters, split into two joins and add the first join's duration_seconds to every offset of the second.

Cost of the whole book

Each concat is a flat $0.01 job, per the docs, and there is no provider inference. The text-to-speech jobs are the real cost. A 3,000-character chapter is 3,000 characters at $47.50 per million, 14.25 cents, billed as 15 cents. Eight such chapters cost $1.20 in TTS plus $0.01 for the join. The 1,200-second cap on a single TTS job is the other reason to work chapter by chapter; the audiobook chapters post covers how to pick the break points.

Format and cap checks

  • Keep WAV for all parts if you will join the result again; MP3 adds priming padding at each edge on every re-encode, according to the docs.
  • Parts need the same channel layout, or the job fails with audio_parts_channel_mismatch.
  • The joined output is capped at 1800 seconds, so a three-hour book is several joined files, not one.

Turning offsets into a chapter list

Each entry in segments[] has an index, a start and a duration_seconds, so the chapter list is the starts in order. Format each start as minutes and seconds, label it with the title of the part at the same index, and the first chapter begins at 0.

Keep your own list of titles in the same order as the parts[] you sent, since the segments are matched by position. If you reorder parts, reorder the titles before you build the list.

A platform that reads chapters from a description usually expects the first stamp at 0:00 and each one on its own line. Check the format of the platform you publish to before you rely on a layout.

  • Chapter start = segment start.
  • Titles follow the order of parts[].
  • First stamp at 0:00.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume