How to create an SRT file from text: time it with TTS

An SRT file needs a start and end time for every line. Voice your text with TTS word timings, then write each timed sentence as a numbered block.

6 min readSume
All posts

To create an SRT file from text, give every line of the text a start and end time, then write the lines as numbered blocks: the number, a start --> end timing line, and the text. For narration, those times have to come from the speech itself, so the text needs audio first. If an AI voice will read it, generate the voice with word and sentence timings and take each block's times from them; they match the narration because they come from the same synthesis.

With Sume, text to speech with timestamps: { "words": true } and segmentation: { "mode": "sentence" } returns gapless segments[]; in current code each one carries its sentence's index, text, start, and end in seconds. The fields come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28; result details called current behavior are read from Sume's code. The SRT layout is the plain-text subtitle format, not a Sume feature.

Can I make an SRT file from text without audio?

Only by guessing. A tool that turns bare text into an SRT has to give each line a made-up duration, for example from its length, and nothing makes a voice read at those times. That can do for on-screen text nobody reads aloud. For narration, time the text from the audio: generate the voice first, as below, or transcribe the recording if one exists.

How do I get timings for my text?

Send the text to POST /v1/tts-1.0/generate with both options, poll the job, and read segments[] and words[] from the result:

  • segmentation requires timestamps.words: true, and "sentence" is its only mode.
  • Segments are gapless: each one ends where the next begins. The cut sits boundary_lead_ms after a sentence's last word (0–500, default 70), and the next segment absorbs the pause.
  • MP3 output is fine. With mp3 you get the timings without per-sentence audio files; those need wav or raw.
  • The same job produces the narration, so the timings belong to that exact audio file.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: course-intro-srt-001" \
  -d '{
    "transcript": "Welcome to the course. Today we set up the project. Then we write a first test.",
    "avatar_handle": "course_host",
    "language": "en",
    "timestamps": { "words": true },
    "segmentation": { "mode": "sentence" }
  }'

How do I turn the segments into an SRT file?

Write one block per segment. Take the times from the segment, and the text from your own script: in current code a segment's text is rebuilt from the engine's word tokens, joined with spaces only when a token holds Latin letters or digits, so Korean text comes back without its word spaces.

Current code ends a sentence at a word ending in ., !, ?, 。, !, ?, or …, even when closing quote marks, ), or ] follow it, so split your script at the same marks and check that both lists have the same length.

SRT parts mapped to TTS result fields, from the TTS schema in the Sume API reference and Sume's current code, read 2026-09-28.
SRT partTake it fromNote
Block numberSegment index + 1index starts at 0
Start timeSegment startIncludes the pause before the sentence; the first segment starts at 0
End timeSegment endEquals the next segment's start
TextYour script's sentence, in orderSegment text is rebuilt from word tokens
def timecode(t):
    ms = round(t * 1000)
    h, ms = divmod(ms, 3_600_000)
    m, ms = divmod(ms, 60_000)
    s, ms = divmod(ms, 1000)
    return f"{h:02}:{m:02}:{s:02},{ms:03}"

def tts_to_srt(result, sentences):
    segments = result["segments"]
    if len(segments) != len(sentences):
        raise ValueError("split the script at the same end marks")
    blocks = []
    for seg, text in zip(segments, sentences):
        times = f"{timecode(seg['start'])} --> {timecode(seg['end'])}"
        blocks.append(f"{seg['index'] + 1}\n{times}\n{text}\n")
    return "\n".join(blocks)

Can the subtitles appear only while each line is spoken?

Yes, with the word timings. Segments are gapless and the next segment absorbs the pause, so a segment-timed line appears during the silence before it is spoken, and the first one appears at 0. For tighter cues, start each block at the start of its first word and end it at the end of its last word, taking the words from words[] whose start falls inside the segment. The same list lets you split a long sentence into two shorter cues at a comma. Both are formatting choices, not rules.

What if my text is already recorded?

Then transcribe the recording instead of generating a voice. Speech to text, POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and the same sentence segmentation, returns segments[] with index, text, start, and end; how to generate an SRT file from a video walks through it. To burn subtitles into a picture, note that Sume's captions API does not take SRT uploads: send the lines as cues, as in how to burn an SRT file into a video.

What does it cost, and what are the limits?

The TTS job bills on the transcript's characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included; in current code the estimate counts characters only, so the timings add nothing. One request takes up to 20,000 characters, and audio longer than 1,200 seconds fails with tts_duration_exceeded, with no credits captured. For longer text, synthesize it in parts; when you join them, shift each part's times by the total length of the parts before it.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume