TTS word timestamps to Timeline slide starts in Python

Call the Sume TTS Router with word timestamps, find the word that opens each slide, and build the Timeline video array of start times in a short Python script.

4 min readSume
All posts

Ask the TTS Router for timestamps: {words: true}, then match the first word of each slide's line to the returned words[] to get a start time for every slide. Feed those starts into the video[] array of a Timeline 1.0 render. This keeps pictures changing exactly when the narrator reaches them, instead of on a fixed interval.

Script

The function below turns a words list into slide entries. It does not call the network, so you can test it first. Each slide keeps its own start and runs until the next one; the last runs to the audio end.

def slides(words, firsts, urls, total):
    """words: [{word,start,end}], firsts: first word of each slide."""
    norm = [w["word"].lower().strip(".,!?") for w in words]
    starts, i = [], 0
    for first in firsts:
        while i < len(norm) and norm[i] != first.lower():
            i += 1
        if i == len(norm):
            raise ValueError(f"word not found: {first}")
        starts.append(words[i]["start"])
    ends = starts[1:] + [total]
    return [{"source_url": u, "start": s, "duration": round(e - s, 2)}
            for u, s, e in zip(urls, starts, ends)]

words = [{"word": "Welcome", "start": 0.0, "end": 0.4},
         {"word": "Kitchen", "start": 3.1, "end": 3.6}]
print(slides(words, ["Welcome", "Kitchen"], ["a.jpg", "b.jpg"], 6.0))

Request fields

Fields from the Sume OpenAPI contract for POST /v1/tts-router/generate and the pricing page, read 2026-10-01.
FieldMeaning
timestamps.wordstrue returns words[] on the completed job result, each with word, start and end
segmentationmode sentence needs timestamps.words and returns gapless segments
Price$0.0475 per 1,000 characters
TranscriptUp to 20000 characters

Limits

A repeated first word matches the first occurrence after the previous slide, so pick distinctive opening words. Timestamps describe the audio you generated, not a different edit of it. In Timeline, video[0].start must be 0, later starts must increase, each slot must be at least 0.2 seconds long and coverage may trail the spine by at most 0.5 seconds, so run the unbilled plan call before rendering. Slide URLs must be your workspace's media.sume.com files, and the voice also needs a voice.id or avatar selector on the TTS request.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume