Text to speech time calculator: estimate, then measure

Estimate speech time as word count divided by a speaking rate, then measure the real file: Sume's TTS reports its duration with word timestamps.

5 min readSume
All posts

To calculate how long a text will take as speech, divide its word count by a speaking rate. If a voice speaks 150 words per minute, for example, a 300-word script takes 2 minutes. The rate changes with the voice and its speed setting, so treat that figure as an estimate and then measure the real audio. With Sume's text to speech, a request with word timestamps returns the file's exact length as duration_seconds in current code.

The Sume facts come from the TTS schema in the Sume API reference, the OpenAPI document behind the API reference docs, read on 2026-09-28. Result fields and checks called current behavior are read from Sume's code, and the price from the code behind API pricing. The formula itself is plain arithmetic.

How do I estimate speech time from a word count?

Seconds equal words divided by words per second; minutes equal words divided by words per minute. The rate is the part to get right. Measure it from a sample read by the same voice at the same speed, because the voice, its speed setting, and the text itself all move it. Sume's avatar-video planner, for instance, estimates 2.8 words per second in current code; that is its estimate for avatar clips, not a measured rate for any TTS voice (how many words fit in 60 seconds).

For a language written without spaces between words, such as Japanese, count characters instead and measure a characters-per-second rate the same way.

// Count words the same way for the estimate and the calibration
function countWords(text) {
  return text.trim().split(/\s+/).filter(Boolean).length;
}

// Calibrate: result = data.result of a finished TTS job that voiced
// sampleText with timestamps: { words: true }
function wordsPerSecond(sampleText, result) {
  return countWords(sampleText) / result.duration_seconds;
}

// Estimate a new script with the rate measured on the same voice and speed
function estimateSeconds(script, rate) {
  return countWords(script) / rate;
}

How do I measure the exact length of the speech?

Generate it with timestamps: { "words": true } on POST /v1/tts-1.0/generate. The completed result then carries words[], each word with its start and end in seconds; in current code it also carries duration_seconds, the length of the audio file. The last word's end marks when the speech stops; duration_seconds runs to the end of the file. Calculate speech rate in words per minute does the same arithmetic for a recording.

What a TTS result tells you about length, from the TTS schema in the Sume API reference and Sume's current code, read 2026-09-28.
FieldWhen it's thereWhat it measures
duration_secondsWith timestamps.words: true (current code)The audio file's length in seconds
words[] start / endWith timestamps.words: trueEach word's timing in seconds; the last end is when speech stops
segments[]segmentation: { "mode": "sentence" }, which needs timestamps.wordsEach sentence's start, end, and duration_seconds (current code)
character_countEvery result (current code)The transcript's length in characters, the unit TTS bills on

How do I make a voiceover fit 15, 30, or 60 seconds?

Measure the first take, then close the gap with words or pace, and measure again.

  • Change the words. At a rate measured on the same voice and speed, you know roughly how many words the gap is worth.
  • Change the pace. generation_config.speed is a speed multiplier from 0.6 to 1.5. The schema doesn't say how a multiplier maps to duration, so read duration_seconds again after every change.
  • Skip the old speed enum (slow, normal, fast): the schema marks it deprecated in favor of generation_config.speed.
  • For a voiceover under a video cut to the second, how to make a 30 second advertisement with AI sets the render's length to match.

How long can one text to speech file be?

One request takes up to 20,000 characters of transcript, and synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, with no credits captured. For a script that runs longer, split it and synthesize each part, as in text to speech for long text.

What does measuring with text to speech cost?

Each measuring run is a TTS job billed on its transcript characters at $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, spaces and punctuation included. In current code the cost estimate counts characters only, so asking for word timestamps adds nothing to it. Calibrate on one short paragraph, reuse that rate for estimates, and synthesize the full script once it should fit.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume