Sume TTS speed: generation_config 0.6-1.5, a slow/normal/fast enum too

Sume TTS has two speed inputs: a speed enum (slow, normal, fast) and generation_config.speed from 0.6 to 1.5. Volume runs 0.5 to 2. Use one, then listen.

4 min readSume
All posts

Sume's TTS request has two places to set speed: a top-level speed field that accepts slow, normal or fast, and generation_config.speed, a number from 0.6 to 1.5. Volume is a number from 0.5 to 2, and emotion is a free string of 1 to 64 characters. The public docs I checked do not say how the two speed fields combine, so send one of them and test.

The ranges

generation_config is a strict object, so unknown keys fail validation. All three of its keys are optional.

Sume rates and limits, read 2026-10-06 from the repo catalog and schemas
FieldTypeRange
speed (top level)Enumslow, normal, fast
generation_config.speedNumber0.6 to 1.5
generation_config.volumeNumber0.5 to 2
generation_config.emotionString1 to 64 characters
{
  "voice": {"id": "<voice uuid>"},
  "transcript": "Welcome back.",
  "generation_config": {"speed": 1.1, "volume": 1.0}
}

Which to use

Use generation_config.speed when you need a number you can reproduce and write down. A series that always runs at 1.1 stays at 1.1. Use the enum for quick drafts where slow, normal or fast is all you need. Mixing them makes it hard to say which one you tested.

A small test plan

Take one 300-character passage. Render it at speed 0.9, 1.0 and 1.1 with the same voice, and note the finished length of each. Three takes of 300 characters cost 2 cents each, so the test is 6 cents. Pick the pace that fits your slot, then lock it in your series settings.

Do the same for volume if you mix with music: render at 1.0 and compare the loudness against your music bed, rather than changing the number blindly.

Record what you used

Whatever you pick, write the exact request settings next to the finished file: speed, volume, emotion, language, voice id and model. A voiceover that has to be redone in six months, for one changed sentence, should come out at the same pace as the rest. Without the record, you are matching by ear.

Fitting a slot

Speed is also how you fit narration into a fixed length, but the ceiling is 1.5. If a translated line needs more than about that, no setting rescues it, so shorten the text. Measure the finished audio and adjust, rather than computing from the character count, because pauses and punctuation change timing. Save the settings with the job result so the next episode starts from the same numbers.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume