Sume TTS speed 0.6 to 1.5: a 14-minute script and the 1,200 s cap
generation_config.speed runs 0.6 to 1.5. A 14-minute read slowed to 0.6 would need about 23 minutes and hit the 1,200-second cap; the price stays per character.

On Sume TTS the speed control is generation_config.speed, a number from 0.6 to 1.5, and it does not change the price, because TTS is billed per character. It can change whether a job succeeds: synthesized audio over 1,200 seconds fails with tts_duration_exceeded. If speed scales duration inversely (our assumption; check with a test line), a 14-minute read at speed 0.6 would run about 23.3 minutes and fail.
The ranges are from the TTS request schema in the Sume repository and the catalog (read 2026-10-09). The same object takes volume (0.5 to 2.0) and an emotion string up to 64 characters. The older speed enum (slow, normal, fast) is marked deprecated.
What the speed range does to a 14-minute script
Take a script that reads in 14 minutes, or 840 seconds, at speed 1.0. Under the inverse assumption, duration = 840 / speed. At 0.6 that is 1,400 seconds, 200 over the cap. At 0.7 it is exactly 1,200 seconds, right at the cap, which leaves no margin. At 0.8 it is 1,050 seconds. At 1.5 it is 560 seconds.
The price is the same in all of those: it depends on the transcript characters, not the speed. A 12,000-character script costs 12 x $0.0475 = $0.57 at every setting.
| speed | Duration (seconds) | Against the 1,200 s cap |
|---|---|---|
| 0.6 | 1,400 | Over: fails |
| 0.7 | 1,200 | At the cap: no margin |
| 0.8 | 1,050 | Under |
| 1.0 | 840 | Under |
| 1.5 | 560 | Under |
Reading back what the job used
A completed text-to-speech job records model_id, voice, language, output_format, generation_config and speed, and each is null when the request did not send it. Sume's jobs documentation says to read those values from the job to make the next line sound the same. For a long script split into parts, send the same generation_config on every part so the pace does not drift between files.
Do not combine the deprecated speed enum with generation_config.speed; pick the numeric field.
A safe workflow for slow reads
Run one test line at the speed you want, measure the audio seconds against the characters you sent, then divide your real script into parts that stay under 1,200 seconds with margin. The jobs and results guide covers polling and the same-key retry if a submit times out.
Volume and emotion
volume is a multiplier from 0.5 to 2.0. It changes loudness only, so it cannot cause a duration failure. emotion is a free string of 1 to 64 characters that guides the delivery; the schema does not list accepted values, so test the words you plan to use on a short line and compare the results.
A practical default for a long read is to set speed and volume once in a shared config object and reuse it on every part. That keeps the pace and level consistent, and it makes the job records comparable.
Source of truth for the limits: the request schema and the catalog, read on 2026-10-09. If a limit changes, the schema wins over this post, so re-read it before you build a long-read pipeline around the 0.7 edge case.
Sources
Related posts
More in Developers
- Sume TTS: 20,000 characters or 1,200 seconds, which fails first?
At 15 characters a second, 20,000 characters is 1,333 seconds, so the 1,200-second audio cap trips first. Split near 15,000 characters per job.
- Sume TTS default is mp3: sentence slices need wav and emit_audio
Sume TTS defaults to mp3 at 44.1 kHz and 128 kbps. Per-sentence audio slices need wav (pcm_s16le) plus segmentation emit_audio, as in the Python request below.
- A 5-minute voice-over: 4.8 MB as 128 kbps MP3, 26.5 MB as WAV
Sume TTS defaults to MP3 at 44,100 Hz and 128 kbps. Five minutes is 4.8 MB; 16-bit mono WAV is 26.46 MB at 44.1 kHz and 28.8 MB at 48 kHz.
- TTS sentence slices: segmentation, wav output and boundary_lead_ms 70
Sume TTS can return gapless sentence segments. Needs timestamps.words and wav or raw for per-sentence audio_url. A 900-character script costs $0.04275.
Written by Sume