How to create an SRT file from text: time it with TTS
An SRT file needs a start and end time for every line. Voice your text with TTS word timings, then write each timed sentence as a numbered block.

To create an SRT file from text, give every line of the text a start and end time, then write the lines as numbered blocks: the number, a start --> end timing line, and the text. For narration, those times have to come from the speech itself, so the text needs audio first. If an AI voice will read it, generate the voice with word and sentence timings and take each block's times from them; they match the narration because they come from the same synthesis.
With Sume, text to speech with timestamps: { "words": true } and segmentation: { "mode": "sentence" } returns gapless segments[]; in current code each one carries its sentence's index, text, start, and end in seconds. The fields come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28; result details called current behavior are read from Sume's code. The SRT layout is the plain-text subtitle format, not a Sume feature.
Can I make an SRT file from text without audio?
Only by guessing. A tool that turns bare text into an SRT has to give each line a made-up duration, for example from its length, and nothing makes a voice read at those times. That can do for on-screen text nobody reads aloud. For narration, time the text from the audio: generate the voice first, as below, or transcribe the recording if one exists.
How do I get timings for my text?
Send the text to POST /v1/tts-1.0/generate with both options, poll the job, and read segments[] and words[] from the result:
segmentationrequirestimestamps.words: true, and"sentence"is its only mode.- Segments are gapless: each one ends where the next begins. The cut sits
boundary_lead_msafter a sentence's last word (0–500, default 70), and the next segment absorbs the pause. - MP3 output is fine. With
mp3you get the timings without per-sentence audio files; those needwavorraw. - The same job produces the narration, so the timings belong to that exact audio file.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: course-intro-srt-001" \
-d '{
"transcript": "Welcome to the course. Today we set up the project. Then we write a first test.",
"avatar_handle": "course_host",
"language": "en",
"timestamps": { "words": true },
"segmentation": { "mode": "sentence" }
}'How do I turn the segments into an SRT file?
Write one block per segment. Take the times from the segment, and the text from your own script: in current code a segment's text is rebuilt from the engine's word tokens, joined with spaces only when a token holds Latin letters or digits, so Korean text comes back without its word spaces.
Current code ends a sentence at a word ending in ., !, ?, 。, !, ?, or …, even when closing quote marks, ), or ] follow it, so split your script at the same marks and check that both lists have the same length.
| SRT part | Take it from | Note |
|---|---|---|
| Block number | Segment index + 1 | index starts at 0 |
| Start time | Segment start | Includes the pause before the sentence; the first segment starts at 0 |
| End time | Segment end | Equals the next segment's start |
| Text | Your script's sentence, in order | Segment text is rebuilt from word tokens |
def timecode(t):
ms = round(t * 1000)
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
def tts_to_srt(result, sentences):
segments = result["segments"]
if len(segments) != len(sentences):
raise ValueError("split the script at the same end marks")
blocks = []
for seg, text in zip(segments, sentences):
times = f"{timecode(seg['start'])} --> {timecode(seg['end'])}"
blocks.append(f"{seg['index'] + 1}\n{times}\n{text}\n")
return "\n".join(blocks)Can the subtitles appear only while each line is spoken?
Yes, with the word timings. Segments are gapless and the next segment absorbs the pause, so a segment-timed line appears during the silence before it is spoken, and the first one appears at 0. For tighter cues, start each block at the start of its first word and end it at the end of its last word, taking the words from words[] whose start falls inside the segment. The same list lets you split a long sentence into two shorter cues at a comma. Both are formatting choices, not rules.
What if my text is already recorded?
Then transcribe the recording instead of generating a voice. Speech to text, POST /v1/stt-1.0/transcribe with a public HTTPS audio_url and the same sentence segmentation, returns segments[] with index, text, start, and end; how to generate an SRT file from a video walks through it. To burn subtitles into a picture, note that Sume's captions API does not take SRT uploads: send the lines as cues, as in how to burn an SRT file into a video.
What does it cost, and what are the limits?
The TTS job bills on the transcript's characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included; in current code the estimate counts characters only, so the timings add nothing. One request takes up to 20,000 characters, and audio longer than 1,200 seconds fails with tts_duration_exceeded, with no credits captured. For longer text, synthesize it in parts; when you join them, shift each part's times by the total length of the parts before it.
Sources
Related posts
More in Developers
- Creatify API: AI avatar videos, Aurora lip sync and credits
The Creatify API makes avatar videos from text or audio, Aurora talking videos from a photo and audio, and ads from a URL, billed in plan credits.
- Cron expression with 6 fields: seconds first or year last?
Standard cron has 5 fields. A 6-field cron expression adds seconds at the front (Spring, Quartz, Azure Functions) or a year at the end (AWS).
- curl bearer token: how to send the Authorization header
Send a bearer token with curl in an Authorization: Bearer header, in double quotes so $TOKEN expands, or use --oauth2-bearer. Fixes for each 401.
- Do you need a GPU for AI video? Only to run the model
You need a GPU for AI video only if you run the model yourself. A hosted video API needs none: your code sends HTTPS and downloads an MP4.
Written by Sume