Sume TTS emotion is a 64 character string: write a guide that fits
The emotion field in Sume TTS generation_config takes 1 to 64 characters. How to write a short, usable guide and test it against a neutral take.

Keep the emotion guide on a Sume TTS request short: the emotion field in generation_config accepts a string of 1 to 64 characters. A longer string is rejected by validation, so a paragraph of acting notes will not fit. Write two or three plain words about tone and pace, such as a warm, unhurried read, and test it against a take with no emotion set.
Short is not a handicap. A guide that names one clear direction is easier to judge than a pile of adjectives.
What does the spec allow?
The Sume OpenAPI spec describes generation_config as optional volume, speed and emotion controls. emotion is a string with minLength 1 and maxLength 64, described as an optional emotion guide for generation. The object does not allow other keys.
| Field | Type | Limit |
|---|---|---|
| emotion | string | 1 to 64 characters |
| speed | number | 0.6 to 1.5 |
| volume | number | 0.5 to 2.0 |
| Any other key | Not allowed | additionalProperties false |
How do you write a guide that fits?
Describe how it should sound, not what it means. The model reads your transcript for meaning already.
- One direction per request: calm, upbeat or serious.
- Add pace in words only if you are not also setting
speed. - Do not paste character backstory; it will not fit and will not help.
- Do not send an empty string; the minimum length is 1.
How do you test one?
Render the same line twice, with and without the guide, and compare. The check below also refuses a guide that is too long before it spends anything.
import os, requests
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
emotion = os.environ.get("EMOTION", "warm, unhurried")
assert 1 <= len(emotion) <= 64, "emotion must be 1 to 64 characters"
body = {
"transcript": "Welcome. Let's get started.",
"voice": {"id": os.environ["VOICE_ID"]},
"language": "en",
"generation_config": {"emotion": emotion},
}
r = requests.post("https://api.sume.com/v1/tts-1.0/generate",
headers=H, json=body, timeout=60)
r.raise_for_status()
print(r.json()["data"]["job"]["id"])
What else shapes delivery?
Speed and volume are numeric and apply to the whole request. For porting tagged scripts from another vendor, see Eleven v4 inline tags vs Sume's emotion guide, and for talking video see avatar emotion and speaking speed.
Sources
Related posts
More in Developers
- Sume TTS takes transcript or transcript_source, never both
A Sume TTS request accepts exactly one text input: literal transcript, or a transcript_source that points at a stored script. How to pick.
- Sume TTS volume 0.5 to 2.0: set the voiceover level before the mix
generation_config volume is a multiplier from 0.5 to 2.0 on Sume TTS. Use it to match narration loudness across jobs before you join or mix them.
- Check for a newer Sume CLI release in CI without auto-upgrading
sume update --check reports whether a newer GitHub Release exists and changes no files. Run it on a schedule, log it, and bump your pinned tag by pull request.
- Sume uploadFile with raw bytes needs a content type, or no request
uploadFile refuses a Uint8Array or ArrayBuffer without contentType and throws SumeUploadError at step create before any call. Pass a typed Blob or a MIME.
Written by Sume