Text to speech API with emotion: audition four values, keep the winner
Sume TTS 1.0 takes a free-text emotion up to 64 characters in generation_config. Run a four-take audition for 4 cents and store the winner.

People who search for a text to speech API with emotion expect a dropdown of moods. On Sume TTS 1.0 it is a text field. generation_config.emotion accepts a free string of 1 to 64 characters, next to speed (0.6 to 1.5) and volume (0.5 to 2). Because the field is open-ended, no list on a page can tell you which values your chosen voice answers well. You find out by listening.
What the vendor says about it
Cartesia's Sonic 3.6 page (read 2026-10-05) says the model varies pacing and intonation to match the emotional context of the transcript, without SSML tags or explicit instructions. Two consequences follow. First, the words you write carry mood on their own, so an apology that reads like an apology needs less steering. Second, Sume's TTS 1.0 has no SSML field, so do not plan on markup. The emotion value is a nudge on top of the text, not a replacement for good copy.
Run an audition grid
Take one real sentence from the script, about 120 characters, and render it with four values. Four jobs at that length each quote 1 cent after rounding up, so the grid is 4 cents. Use the same voice and an idempotency key per value, so a rerun costs nothing.
import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = {**H, **({"Idempotency-Key": key} if key else {})}
r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
r.raise_for_status()
job = r.json()["data"]["job"]["id"]
while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
time.sleep(3)
res = requests.get(f"{B}/jobs/{job}/result", headers=H)
res.raise_for_status()
return res.json()["data"]["result"]
LINE = "We are sorry the parcel arrived late. Here is a refund, and a code for next time."
for emo in ["apologetic", "warm", "calm", "upbeat"]:
res = run("/tts-1.0/generate", {"transcript": LINE, "language": "en",
"voice": {"id": os.environ["VOICE_ID"]},
"generation_config": {"emotion": emo}}, key="aud-" + emo)
print(emo, res["audio_url"], res["generation_config"])Judge, then record
Listen blind if you can: have a colleague pick without seeing the labels. Then store the winner's generation_config with the voice id, the way the result-logging post describes. The job result echoes the config that the audio was made with, so the log can come from the result rather than from your memory of what you sent.
Rules of thumb that save credits
- Audition on one sentence, not a full script. The full 20,000-character maximum would cost 95 cents per take.
- Change one thing per take. If you also change
speed, you cannot tell which setting you heard. - Keep the value short and concrete. Sixty-four characters is a limit, not a target.
- If a value does nothing audible, the voice may not respond to it. Try a different voice instead of a longer string.
Where it stops
Sume does not return a score for how emotional a take is, and it cannot promise that two voices interpret the same word identically. Treat the grid as listening time, not a benchmark.
Sources
Related posts
More in Developers
- Text-to-video API in Node: submit, poll and download with a deadline
A runnable Node 18+ script that submits a text-to-video job to Sume, polls it with a deadline, and saves the MP4. No dependencies and no top-level await.
- Threads image post: 8 MB, 320 to 1440 px wide, from video frames
Threads images must be JPEG or PNG, up to 8 MB, 320 to 1440 px wide. Pull a still from a Sume clip with video-frames and clamp the long edge to 1440.
- Threads video aspect ratio up to 10:1 and 1920 px: timeline output
Threads allows video ratios from 0.01:1 to 10:1 and 1920 px, and recommends 9:16. The Sume timeline default 1080x1920 fits, and width and height are settable.
- Threads video frame rate 23 to 60 fps: conform with video-trim
Threads accepts 23 to 60 fps video. Sume video-trim output.fps takes 24, 25, 30 or 60, all inside that range. Conform a mixed-rate clip in one call.
Written by Sume