Text to speech time calculator: estimate, then measure
Estimate speech time as word count divided by a speaking rate, then measure the real file: Sume's TTS reports its duration with word timestamps.

To calculate how long a text will take as speech, divide its word count by a speaking rate. If a voice speaks 150 words per minute, for example, a 300-word script takes 2 minutes. The rate changes with the voice and its speed setting, so treat that figure as an estimate and then measure the real audio. With Sume's text to speech, a request with word timestamps returns the file's exact length as duration_seconds in current code.
The Sume facts come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28. Result fields and checks called current behavior are read from Sume's code, and the price from the code behind API pricing. The formula itself is plain arithmetic.
How do I estimate speech time from a word count?
Seconds equal words divided by words per second; minutes equal words divided by words per minute. The rate is the part to get right. Measure it from a sample read by the same voice at the same speed, because the voice, its speed setting, and the text itself all move it. Sume's avatar-video planner, for instance, estimates 2.8 words per second in current code; that is its estimate for avatar clips, not a measured rate for any TTS voice (how many words fit in 60 seconds).
For a language written without spaces between words, such as Japanese, count characters instead and measure a characters-per-second rate the same way.
// Count words the same way for the estimate and the calibration
function countWords(text) {
return text.trim().split(/\s+/).filter(Boolean).length;
}
// Calibrate: result = data.result of a finished TTS job that voiced
// sampleText with timestamps: { words: true }
function wordsPerSecond(sampleText, result) {
return countWords(sampleText) / result.duration_seconds;
}
// Estimate a new script with the rate measured on the same voice and speed
function estimateSeconds(script, rate) {
return countWords(script) / rate;
}How do I measure the exact length of the speech?
Generate it with timestamps: { "words": true } on POST /v1/tts-1.0/generate. The completed result then carries words[], each word with its start and end in seconds; in current code it also carries duration_seconds, the length of the audio file. The last word's end marks when the speech stops; duration_seconds runs to the end of the file. Calculate speech rate in words per minute does the same arithmetic for a recording.
| Field | When it's there | What it measures |
|---|---|---|
duration_seconds | With timestamps.words: true (current code) | The audio file's length in seconds |
words[] start / end | With timestamps.words: true | Each word's timing in seconds; the last end is when speech stops |
segments[] | segmentation: { "mode": "sentence" }, which needs timestamps.words | Each sentence's start, end, and duration_seconds (current code) |
character_count | Every result (current code) | The transcript's length in characters, the unit TTS bills on |
How do I make a voiceover fit 15, 30, or 60 seconds?
Measure the first take, then close the gap with words or pace, and measure again.
- Change the words. At a rate measured on the same voice and speed, you know roughly how many words the gap is worth.
- Change the pace.
generation_config.speedis a speed multiplier from 0.6 to 1.5. The schema doesn't say how a multiplier maps to duration, so readduration_secondsagain after every change. - Skip the old
speedenum (slow,normal,fast): the schema marks it deprecated in favor ofgeneration_config.speed. - For a voiceover under a video cut to the second, how to make a 30 second advertisement with AI sets the render's length to match.
How long can one text to speech file be?
One request takes up to 20,000 characters of transcript, and synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, with no credits captured. For a script that runs longer, split it and synthesize each part, as in text to speech for long text.
What does measuring with text to speech cost?
Each measuring run is a TTS job billed on its transcript characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included. In current code the cost estimate counts characters only, so asking for word timestamps adds nothing to it. Calibrate on one short paragraph, reuse that rate for estimates, and synthesize the full script once it should fit.
Sources
Related posts
More in Use cases
- TikTok Commercial Music Library (CML): what it covers
TikTok's Commercial Music Library is a pre-cleared set of songs businesses may use free, but only on TikTok. What it covers, and music for other uses.
- Transcribe a lecture to notes with timestamps
Transcribe a lecture with sentence timestamps, then have a language model turn it into notes whose headings point back to where each topic starts.
- Video prospecting with an AI avatar: one clip per prospect
Video prospecting puts a short personal video in a sales outreach message. With an AI avatar, fill one script template per prospect and render each clip.
- What is a video sales letter (VSL)? And making one with AI
A video sales letter (VSL) is a sales pitch delivered as one narrated video, from hook to offer. What goes in one, and how to build it with AI parts.
Written by Sume