Text to speech API with emotion: audition four values, keep the winner

Sume TTS 1.0 takes a free-text emotion up to 64 characters in generation_config. Run a four-take audition for 4 cents and store the winner.

4 min readSume
All posts

People who search for a text to speech API with emotion expect a dropdown of moods. On Sume TTS 1.0 it is a text field. generation_config.emotion accepts a free string of 1 to 64 characters, next to speed (0.6 to 1.5) and volume (0.5 to 2). Because the field is open-ended, no list on a page can tell you which values your chosen voice answers well. You find out by listening.

What the vendor says about it

Cartesia's Sonic 3.6 page (read 2026-10-05) says the model varies pacing and intonation to match the emotional context of the transcript, without SSML tags or explicit instructions. Two consequences follow. First, the words you write carry mood on their own, so an apology that reads like an apology needs less steering. Second, Sume's TTS 1.0 has no SSML field, so do not plan on markup. The emotion value is a nudge on top of the text, not a replacement for good copy.

Run an audition grid

Take one real sentence from the script, about 120 characters, and render it with four values. Four jobs at that length each quote 1 cent after rounding up, so the grid is 4 cents. Use the same voice and an idempotency key per value, so a rerun costs nothing.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

LINE = "We are sorry the parcel arrived late. Here is a refund, and a code for next time."
for emo in ["apologetic", "warm", "calm", "upbeat"]:
    res = run("/tts-1.0/generate", {"transcript": LINE, "language": "en",
              "voice": {"id": os.environ["VOICE_ID"]},
              "generation_config": {"emotion": emo}}, key="aud-" + emo)
    print(emo, res["audio_url"], res["generation_config"])

Judge, then record

Listen blind if you can: have a colleague pick without seeing the labels. Then store the winner's generation_config with the voice id, the way the result-logging post describes. The job result echoes the config that the audio was made with, so the log can come from the result rather than from your memory of what you sent.

Rules of thumb that save credits

  • Audition on one sentence, not a full script. The full 20,000-character maximum would cost 95 cents per take.
  • Change one thing per take. If you also change speed, you cannot tell which setting you heard.
  • Keep the value short and concrete. Sixty-four characters is a limit, not a target.
  • If a value does nothing audible, the voice may not respond to it. Try a different voice instead of a longer string.

Where it stops

Sume does not return a score for how emotional a take is, and it cannot promise that two voices interpret the same word identically. Treat the grid as listening time, not a benchmark.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume