Text to speech API with emotion, speed and volume controls
Sume's text to speech API takes generation_config: speed 0.6 to 1.5, volume 0.5 to 2.0 and an optional emotion guide of up to 64 characters.

Yes, Sume's text to speech API has three delivery controls, all inside one optional generation_config object: speed from 0.6 to 1.5, volume from 0.5 to 2.0, and an emotion guide, a free string of 1 to 64 characters. The same object works on TTS 1.0 and on the TTS Router.
The field limits are from the TTS request schema in the Sume API reference, read 2026-09-29. The one engine fact, about how Sonic 3.6 handles emotion, is from Cartesia's Sonic 3.6 docs.
What can I set in generation_config?
All three fields are optional, and the object rejects any other key. Speed and volume are multipliers.
| Field | Type and range | What the schema says |
|---|---|---|
speed | Number, 0.6 to 1.5 | Speed multiplier |
volume | Number, 0.5 to 2.0 | Volume multiplier |
emotion | String, 1 to 64 characters | Optional emotion guide for generation |
How do I ask for a calmer, slower read?
Put the three fields in generation_config next to your transcript and voice. The emotion value below is a plain word chosen for the example: the schema gives no list of accepted values, so treat it as a guide and listen to the result.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-calm-001" \
-d '{
"transcript": "Take a slow breath. We start in a minute.",
"avatar_handle": "acme",
"language": "en",
"generation_config": { "speed": 0.9, "volume": 1.0, "emotion": "calm" }
}'Does the engine also read emotion from the text?
Cartesia says Sonic 3.6 varies its pacing and intonation to match the emotional context of the transcript, without SSML tags or explicit instructions. Pause length adapts to the sentence, and in-transcript disfluencies such as "uhm" or "hmm" produce a natural thinking pace.
So the words you write carry delivery too. Sume's schema documents no inline tags or SSML: write the line the way it should sound, and use emotion, speed and volume for the rest. Sume's TTS 1.0 route has no engine picker; to name a Sonic version, use the TTS Router; Text to speech API covers the request.
Does it work on the TTS Router too?
Yes. The router request schema carries the same generation_config object, so you can name a Sonic version and still set the three controls. Send model as well; on TTS 1.0 a model field is rejected.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: tts-router-calm-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Take a slow breath. We start in a minute.",
"avatar_handle": "acme",
"language": "en",
"generation_config": { "speed": 0.9, "emotion": "calm" }
}'What are the limits?
- Stay inside 0.6 to 1.5 for speed and 0.5 to 2.0 for volume: those are the schema's minimum and maximum.
emotionhas no documented value list, and the docs make no promise about how strongly it changes the voice. Compare takes before you ship.- Pronunciation problems are a different fix: see Text to speech pronunciation.
- To budget a script's length before you synthesize, see the text to speech time calculator.
Sources
Related posts
More in Models
- Text to speech sample rate: which output_format to pick
Sume's TTS output_format takes mp3, wav or raw, six sample rates from 8000 to 48000 Hz and four PCM encodings. The default, and when to send wav.
- Veo 3 shutdown: which Gemini API models retired and what replaces them
Google retired veo-3.0-generate-001, veo-3.0-fast-generate-001 and veo-2.0-generate-001 on 2026-06-30. Its listed replacements are the Veo 3.1 preview models.
- Veo 3.1 4K API: 8 seconds only, and 4K options on Sume
Veo 3.1 makes 4K video only at 8 seconds, and not on the Lite model. On Sume, gemini-omni-flash-1.1 lists 4K for 3 to 10 seconds.
- Veo 3.1 API: how to call it, and what Sume lists instead
The Veo 3.1 API is Google's, called through the Gemini API with preview model codes. Sume does not list Veo; here are the catalog models with the same inputs.
Written by Sume