Sume TTS speed 0.6 to 1.5: a 14-minute script and the 1,200 s cap

generation_config.speed runs 0.6 to 1.5. A 14-minute read slowed to 0.6 would need about 23 minutes and hit the 1,200-second cap; the price stays per character.

5 min readSume
All posts

On Sume TTS the speed control is generation_config.speed, a number from 0.6 to 1.5, and it does not change the price, because TTS is billed per character. It can change whether a job succeeds: synthesized audio over 1,200 seconds fails with tts_duration_exceeded. If speed scales duration inversely (our assumption; check with a test line), a 14-minute read at speed 0.6 would run about 23.3 minutes and fail.

The ranges are from the TTS request schema in the Sume repository and the catalog (read 2026-10-09). The same object takes volume (0.5 to 2.0) and an emotion string up to 64 characters. The older speed enum (slow, normal, fast) is marked deprecated.

What the speed range does to a 14-minute script

Take a script that reads in 14 minutes, or 840 seconds, at speed 1.0. Under the inverse assumption, duration = 840 / speed. At 0.6 that is 1,400 seconds, 200 over the cap. At 0.7 it is exactly 1,200 seconds, right at the cap, which leaves no margin. At 0.8 it is 1,050 seconds. At 1.5 it is 560 seconds.

The price is the same in all of those: it depends on the transcript characters, not the speed. A 12,000-character script costs 12 x $0.0475 = $0.57 at every setting.

A 840-second read at different speeds, assuming duration scales as 1/speed
speedDuration (seconds)Against the 1,200 s cap
0.61,400Over: fails
0.71,200At the cap: no margin
0.81,050Under
1.0840Under
1.5560Under

Reading back what the job used

A completed text-to-speech job records model_id, voice, language, output_format, generation_config and speed, and each is null when the request did not send it. Sume's jobs documentation says to read those values from the job to make the next line sound the same. For a long script split into parts, send the same generation_config on every part so the pace does not drift between files.

Do not combine the deprecated speed enum with generation_config.speed; pick the numeric field.

A safe workflow for slow reads

Run one test line at the speed you want, measure the audio seconds against the characters you sent, then divide your real script into parts that stay under 1,200 seconds with margin. The jobs and results guide covers polling and the same-key retry if a submit times out.

Volume and emotion

volume is a multiplier from 0.5 to 2.0. It changes loudness only, so it cannot cause a duration failure. emotion is a free string of 1 to 64 characters that guides the delivery; the schema does not list accepted values, so test the words you plan to use on a short line and compare the results.

A practical default for a long read is to set speed and volume once in a shared config object and reuse it on every part. That keeps the pace and level consistent, and it makes the job records comparable.

Source of truth for the limits: the request schema and the catalog, read on 2026-10-09. If a limit changes, the schema wins over this post, so re-read it before you build a long-read pipeline around the 0.7 edge case.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume