Sume TTS speed: generation_config 0.6-1.5, a slow/normal/fast enum too
Sume TTS has two speed inputs: a speed enum (slow, normal, fast) and generation_config.speed from 0.6 to 1.5. Volume runs 0.5 to 2. Use one, then listen.

Sume's TTS request has two places to set speed: a top-level speed field that accepts slow, normal or fast, and generation_config.speed, a number from 0.6 to 1.5. Volume is a number from 0.5 to 2, and emotion is a free string of 1 to 64 characters. The public docs I checked do not say how the two speed fields combine, so send one of them and test.
The ranges
generation_config is a strict object, so unknown keys fail validation. All three of its keys are optional.
| Field | Type | Range |
|---|---|---|
| speed (top level) | Enum | slow, normal, fast |
| generation_config.speed | Number | 0.6 to 1.5 |
| generation_config.volume | Number | 0.5 to 2 |
| generation_config.emotion | String | 1 to 64 characters |
{
"voice": {"id": "<voice uuid>"},
"transcript": "Welcome back.",
"generation_config": {"speed": 1.1, "volume": 1.0}
}Which to use
Use generation_config.speed when you need a number you can reproduce and write down. A series that always runs at 1.1 stays at 1.1. Use the enum for quick drafts where slow, normal or fast is all you need. Mixing them makes it hard to say which one you tested.
A small test plan
Take one 300-character passage. Render it at speed 0.9, 1.0 and 1.1 with the same voice, and note the finished length of each. Three takes of 300 characters cost 2 cents each, so the test is 6 cents. Pick the pace that fits your slot, then lock it in your series settings.
Do the same for volume if you mix with music: render at 1.0 and compare the loudness against your music bed, rather than changing the number blindly.
Record what you used
Whatever you pick, write the exact request settings next to the finished file: speed, volume, emotion, language, voice id and model. A voiceover that has to be redone in six months, for one changed sentence, should come out at the same pace as the rest. Without the record, you are matching by ear.
Fitting a slot
Speed is also how you fit narration into a fixed length, but the ceiling is 1.5. If a translated line needs more than about that, no setting rescues it, so shorten the text. Measure the finished audio and adjust, rather than computing from the character count, because pauses and punctuation change timing. Save the settings with the job result so the next episode starts from the same numbers.
Sources
Related posts
More in Developers
- Sume TTS output_format: the mp3 44.1 kHz 128 kbps default and options
Send no output_format and Sume TTS returns mp3, 44,100 Hz, 128 kbps. Here is every allowed container, sample rate, bit rate and raw encoding.
- Sume TTS rate math: 38 millionths a character to $0.0475 per 1,000
How the TTS Router catalog's 38 micro-dollars per character, times Sume's 1.25 margin, becomes $0.0475 per 1,000 characters, and how to check it yourself.
- Preview Sume TTS sentence ids, lengths and job cost before you submit
A short Python script that splits a script like Sume's source API, groups sentences under 20,000 characters and prices each job at $0.0475 per 1,000 characters.
- Why Sume's video models list shows 11 ids, not 12
Docs describe 12 video router ids, but GET /v1/videos/models can return 11. higgsfield-genjutsu lists only when its provider is configured. Check yours.
Written by Sume