Emotion tags in the script or an emotion field: Sume's bill
MAI-Voice takes emotion tags. Sume's documented control is generation_config.emotion, and every transcript character is billed, tags included.

On Sume, steer delivery with the generation_config.emotion field, not with bracketed tags typed into the transcript. Sume's TTS contract documents the field, and it bills every transcript character, spaces and punctuation included, so a tag you type is a character you pay for and the docs do not define tag syntax.
What Microsoft says about tags
Microsoft's MAI-Voice-2 announcement lists emotion tags such as sad, whispered and excited as granular emotion control (read 2026-10-05). The MAI-Voice-2.1 model page marks emotion control as available on both the Standard and Flash variants (read 2026-10-05). Those pages describe Microsoft's own interface. They do not say how any other engine treats the same tags.
What Sume documents
Sume's generated API contract for TTS has an optional generation_config object with three controls: volume (multiplier from 0.5 to 2.0), speed (multiplier from 0.6 to 1.5) and emotion (an optional emotion guide string for generation). The same contract says the transcript is the text to synthesize, with spaces and punctuation counting toward usage, up to 20,000 characters.
Nothing in that contract or in the TTS Router doc describes inline tag syntax inside the transcript. Treat anything you type into the transcript as text the voice may be asked to say, and put direction in the emotion field.
What a tag would cost
Take a 1,350-character, 90-second script split into 12 lines. A tag such as [excited] is 10 characters, so 12 tags add 120 characters. At Sume's $0.0475 per 1,000 characters, 120 characters is 0.57 cent raw.
| Version | Characters | Raw cost | Billed (rounded up to a cent) |
|---|---|---|---|
| No tags | 1,350 | $0.0641 | $0.07 |
| 12 tags of 10 characters | 1,470 | $0.0698 | $0.07 |
| 2,000-character script plus 12 tags | 2,120 | $0.1007 | $0.11 |
The money is not the problem. The problem is that one emotion string applies to the whole request, so per-line direction means per-line jobs, and every job costs at least one cent. Twelve one-line jobs at about 110 characters each are 12 cents against 7 cents for one job, which is the rounding effect at work.
A practical split
Test the guide wording on a short line first. A take of 80 characters costs 1 cent, so five wordings cost 5 cents, which is cheap next to an ad that ships with the wrong tone.
- One mood across the read: one job, one
emotionguide, onespeed. - Two moods, such as a calm intro and an excited call to action: two jobs, then join them with Sume's timeline audio concat at $0.01 per job.
- Three or more moods: ask whether the script should be three videos' worth of reads, because each extra job adds at least a cent and a seam to check.
Check word timing after you change the tone
A 90-second read at 1,350 characters is about 15 characters a second under the assumption used across these posts. A guide that asks for a slower, warmer delivery can move that to 12, which stretches the same script to roughly 112 seconds. Measure it rather than trusting the arithmetic.
Request timestamps.words: true. A slower, warmer read stretches the script, and the last word's end time tells you whether it still fits its slot. If it overruns, change speed before you rewrite the copy.
A test you can run in ten minutes
Pick one 200-character line and render it three times: once with no guide, once with a tag in the text, once with the emotion field set and the text left clean. Each job is a single cent, so the whole comparison costs three cents. Listen on the speakers your audience will use, not on studio monitors.
Write down which version you chose and why. A team that keeps this note for a week ends up with a house rule, such as "guide in the field, never in the script", which is far cheaper than relitigating it on every ad. Because tag characters in a script count toward the bill, the rule also keeps receipts predictable.
Sources
Related posts
More in Developers
- 'end_image_url requires image_url': a last frame needs a first frame
A last frame without a first frame is refused on both Sume video routes. Add a first frame, or drop the end frame and describe the ending in the prompt.
- An es-MX voice with an es-ES request: Sume compares primary language
Sume's TTS language check compares regional tags by primary language, so es-MX against es-ES passes. What that means for accent, and what 409 you get otherwise.
- Estimate a Sume video clip in Python: list x 1.25, rounded up
A runnable Python script that prices a clip for Wan 3.0, MiniMax H3, H3 Max and Gemini Omni 1.1 Flash with integer micros, so the cents match the bill.
- Add the EU AI icon to a clip with Timeline compose overlay
The EU publishes an AI icon as PNG and SVG. Sume timeline compose with operation overlay puts a hosted still on a hosted video and returns a new MP4 for $0.02.
Written by Sume