TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
Sume TTS marks the slow, normal and fast speed enum deprecated. Send generation_config.speed from 0.6 to 1.5, plus volume 0.5 to 2 and an emotion string.

Use generation_config.speed, a number from 0.6 to 1.5, to change how fast Sume TTS speaks. The older top-level speed field, with values slow, normal or fast, is marked deprecated in the OpenAPI document, which says to prefer generation_config.speed. The same object also takes volume (0.5 to 2.0) and emotion, a string of 1 to 64 characters.
What are the exact ranges?
From the request schema for TTS 1.0 and the TTS Router: generation_config.volume is a multiplier in [0.5, 2.0]; generation_config.speed is a multiplier in [0.6, 1.5]; generation_config.emotion is a string, minimum length 1 and maximum 64. The object does not allow extra keys, so a misspelled field is refused instead of ignored.
Cartesia's Sonic 3.6 changelog entry (read 2026-10-02) says speed and volume controls behave on 3.6 as they do on Sonic 3.5, so a move to the newer model does not need new values.
| Field | Type and range | Status |
|---|---|---|
generation_config.speed | Number, 0.6 to 1.5 | Preferred |
generation_config.volume | Number, 0.5 to 2.0 | Current |
generation_config.emotion | String, 1 to 64 characters | Current |
speed | slow, normal or fast | Deprecated |
Which should I send?
Send the number. It gives you steps between slow and fast, and it is the one the docs point at. Take 1.0 as your baseline, then adjust in small steps, such as 0.9 for a calmer read or 1.1 for a tighter one, and listen. Do not send both fields on one request unless you have a reason.
Your result will tell you what was used. A finished job records generation_config and speed, each null when you did not send it, as described in Jobs and results.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: speed-090-001" \
-d '{
"transcript": "Take a slow breath in, and out.",
"avatar_handle": "@your_avatar",
"language": "en",
"generation_config": { "speed": 0.9, "volume": 1.0, "emotion": "calm" }
}'Does a slower speed change the cost?
TTS usage counts the characters of the transcript, spaces and punctuation included, up to 20000. Speed does not change the character count. It does change the length of the audio, and synthesized audio over 1200 seconds fails with tts_duration_exceeded, so a long script at 0.6 can reach that limit before the character cap.
Sources
Related posts
More in Developers
- TTS call returned processing in sync mode: poll, do not resubmit
Sume TTS in sync mode waits at most 30 seconds. If the job is not terminal, poll status_url, and retry a submit only with the same Idempotency-Key.
- Test call audio for voice agents: Sume TTS at 8 kHz mu-law
Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.
- TTS voice.id: a UUID or a voi_ library id? What Sume accepts
Sume TTS voice.id takes a voice UUID or a voi_ library id. Any other shape fails with 400 invalid_voice_id before a job is queued or credits are reserved.
- TTS word timestamps to Timeline slide starts in Python
Call the Sume TTS Router with word timestamps, find the word that opens each slide, and build the Timeline video array of start times in a short Python script.
Written by Sume