Text to speech API with emotion, speed and volume controls

Sume's text to speech API takes generation_config: speed 0.6 to 1.5, volume 0.5 to 2.0 and an optional emotion guide of up to 64 characters.

4 min readSume
All posts

Yes, Sume's text to speech API has three delivery controls, all inside one optional generation_config object: speed from 0.6 to 1.5, volume from 0.5 to 2.0, and an emotion guide, a free string of 1 to 64 characters. The same object works on TTS 1.0 and on the TTS Router.

The field limits are from the TTS request schema in the Sume API reference, read 2026-09-29. The one engine fact, about how Sonic 3.6 handles emotion, is from Cartesia's Sonic 3.6 docs.

What can I set in generation_config?

All three fields are optional, and the object rejects any other key. Speed and volume are multipliers.

From the TTS request schema in the Sume API reference, read 2026-09-29.
FieldType and rangeWhat the schema says
speedNumber, 0.6 to 1.5Speed multiplier
volumeNumber, 0.5 to 2.0Volume multiplier
emotionString, 1 to 64 charactersOptional emotion guide for generation

How do I ask for a calmer, slower read?

Put the three fields in generation_config next to your transcript and voice. The emotion value below is a plain word chosen for the example: the schema gives no list of accepted values, so treat it as a guide and listen to the result.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-calm-001" \
  -d '{
    "transcript": "Take a slow breath. We start in a minute.",
    "avatar_handle": "acme",
    "language": "en",
    "generation_config": { "speed": 0.9, "volume": 1.0, "emotion": "calm" }
  }'

Does the engine also read emotion from the text?

Cartesia says Sonic 3.6 varies its pacing and intonation to match the emotional context of the transcript, without SSML tags or explicit instructions. Pause length adapts to the sentence, and in-transcript disfluencies such as "uhm" or "hmm" produce a natural thinking pace.

So the words you write carry delivery too. Sume's schema documents no inline tags or SSML: write the line the way it should sound, and use emotion, speed and volume for the rest. Sume's TTS 1.0 route has no engine picker; to name a Sonic version, use the TTS Router; Text to speech API covers the request.

Does it work on the TTS Router too?

Yes. The router request schema carries the same generation_config object, so you can name a Sonic version and still set the three controls. Send model as well; on TTS 1.0 a model field is rejected.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tts-router-calm-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Take a slow breath. We start in a minute.",
    "avatar_handle": "acme",
    "language": "en",
    "generation_config": { "speed": 0.9, "emotion": "calm" }
  }'

What are the limits?

  • Stay inside 0.6 to 1.5 for speed and 0.5 to 2.0 for volume: those are the schema's minimum and maximum.
  • emotion has no documented value list, and the docs make no promise about how strongly it changes the voice. Compare takes before you ship.
  • Pronunciation problems are a different fix: see Text to speech pronunciation.
  • To budget a script's length before you synthesize, see the text to speech time calculator.

Sources

Related posts

More in Models

All Models posts

Written by Sume