TTS speed 1.2 turns a 30-second read into about 25 seconds
Sume TTS accepts generation_config.speed from 0.6 to 1.5. Speed changes length, not the bill: 450 characters cost $0.021375 at any speed. Test the real length.

To fit a 30-second voice-over into 25 seconds, set generation_config.speed to 1.2 on the Sume TTS request. The field accepts values from 0.6 to 1.5, and the price does not change, because TTS is billed per character. If the speed only scales playback, 30 / 1.2 = 25 seconds; confirm with a measurement, since the docs do not promise exact scaling.
The request
The OpenAPI reference lists generation_config with volume (0.5 to 2), speed (0.6 to 1.5) and a free-text emotion guide. The older top-level speed enum (slow, normal, fast) is marked deprecated in favor of generation_config.speed. Select the voice with avatar_handle or voice.id, and set language for any non-English script.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-speed-12-001" \
-d '{
"transcript": "Your 450-character script goes here.",
"avatar_handle": "@brandvoice",
"language": "en",
"generation_config": { "speed": 1.2 },
"mode": "sync",
"wait_timeout_seconds": 30
}'Expected lengths
Start from a 30-second read at normal speed and divide by the multiplier. These are expectations to check against the returned file, not guarantees.
| generation_config.speed | Expected length of a 30 s read | Arithmetic |
|---|---|---|
| 0.6 | 50 s | 30 / 0.6 |
| 0.8 | 37.5 s | 30 / 0.8 |
| 1.0 | 30 s | 30 / 1.0 |
| 1.2 | 25 s | 30 / 1.2 |
| 1.5 | 20 s | 30 / 1.5 |
The bill does not move
Sume TTS is $0.0475 per 1,000 characters (Sume API catalog, read 2026-10-09). A script of 450 characters, which is 30 seconds at an assumed 15 characters a second, is 0.45 x $0.0475 = $0.021375 whether you ask for 0.6 or 1.5. The cost lever is the script, not the speed.
That makes speed the cheap fix for a small overrun and rewriting the script the right fix for a large one. A 30 second read squeezed to 20 seconds at 1.5 is a 50 percent speed-up, and you should listen to it before shipping.
Check the result before you cut
Read the audio from the job result, measure its duration, and only then place it on a timeline. Sume's Timeline renders are billed per output minute and rounded up, so the 25-second clip and the 30-second clip both land in the first billed minute; the saving is for the edit, not the invoice. See the jobs and results docs for polling, and use the word timings option if you need to align captions to the faster read.
Speed or rewrite: a quick rule
Use speed for overruns under about 20 percent, which is the 1.0 to 1.2 band in the table, and cut words for anything larger. Cutting words also lowers the bill: removing 90 characters from a 450 character script saves 0.09 x $0.0475 = $0.004275 per take, small per clip but real across a thousand localized variants.
Keep one preset per series. Store the speed with the voice and language so every episode in a series comes out at the same pace, and re-measure when you change voice, because a different voice can read the same text at a different rate. The OpenAPI reference also offers word timestamps and sentence segmentation, which let you check where a faster read now ends each sentence without opening an audio editor.
Sources
Related posts
More in Developers
- TTS word timings straight into captions: a 45-second ad for 33 cents
Ask Sume TTS for timestamps.words and send them as words on the caption job: no speech-to-text. 700 characters, render and captions come to $0.33.
- 20 Nano Banana 2.1 2K images: $3.00 and one jobs_result read
Twenty Nano Banana 2.1 images at 2K cost $3.00 on Sume. Over hosted MCP, one jobs_result call with 20 job_ids reads them back; failed ids are named.
- 20 Omni Flash 1.1 clips in one jobs_wait wave: $12.50 at 720p
Twenty 5-second Gemini Omni Flash 1.1 clips at 720p cost $12.50 on Sume ($0.625 each) and fit one jobs_wait call of 20 ids. Cost table and the call body.
- TypeScript: check 8:1 against /v1/images/models before you POST
A 20-line TypeScript guard that reads the aspect_ratio values for a Sume image model and refuses a ratio it does not list, instead of a 400 at request time.
Written by Sume