Sume TTS has no SSML field: speed, emotion, pronunciation dictionary

Moving an SSML voice script to Sume TTS? The request takes plain transcript text, so use speed 0.6 to 1.5, volume, emotion and a pronunciation dictionary id.

5 min readSume
All posts

No. Sume's TTS request has no SSML or markup field. It takes plain transcript text (1 to 20,000 characters) and a small set of typed controls: generation_config.speed from 0.6 to 1.5, generation_config.volume from 0.5 to 2.0, a free-text emotion guide, a pronunciation_dict_id, and language. If your script is SSML, strip the tags and move each intent to one of those fields.

New voice models launched on 2026-10-01 at $22 per 1M characters (MAI-Voice-2.1) and $15 per 1M characters (Flash) (Microsoft AI, read 2026-10-05), so scripts written for them will come with markup on top of the words.

Where each intent goes

Each markup intent has a Sume field, a coarser one, or none. The table says which. It does not claim that the voices behave the same.

Voice intents on Sume TTS (read 2026-10-05)
Intent in a markup scriptSume TTS fieldRange or form
Faster or slower deliverygeneration_config.speed0.6 to 1.5
Louder or quietergeneration_config.volume0.5 to 2.0
Mood of the readgeneration_config.emotionFree-text guide
Name or term pronunciationpronunciation_dict_idOne dictionary id
Language of the voicelanguageBCP-47 or ISO-639, such as ko or ja
Exact pause lengthNone documentedNot a field

Strip tags before you pay

Characters count whatever you send: the schema says spaces and punctuation count toward usage. Tags left inside the transcript would be billed as text and might be spoken, so remove them before you submit. On the MCP tool, dry_run previews the cost and max_spend_usd caps it.

Language and the older speed field

Set language for every non-English transcript. When it is omitted the provider defaults to English, with a fallback guess for Hangul-only or kana-only text. If the voice and language disagree, Sume returns tts_voice_language_warning, and you retry with confirm_language_mismatch: true only after the user agrees.

The old speed enum with slow, normal and fast is deprecated. Prefer the numeric field.

A request body

A minimal request body, with the controls in one place:

{
  "idempotency_key": "ad-read-001",
  "payload": {
    "transcript": "Spring sale starts Friday. Everything is half price.",
    "language": "en",
    "voice": {"id": "voi_00000000000000000000000000000000"},
    "generation_config": {"speed": 1.1, "volume": 1.0, "emotion": "warm"},
    "output_format": {"container": "mp3", "sample_rate": 44100, "bit_rate": 128000}
  }
}

Notes

That is the MCP shape, with fields inside payload. The HTTP route takes the same fields at the top level. The voice id above is a placeholder, so copy a real id verbatim from your voice library or resolve one from an avatar handle.

Pauses are the gap. If a script depends on exact break lengths, write the pause into the text with punctuation and listen, because the schema offers no timing field for it.

The same goes for emphasis on a single word. The documented controls apply to the whole request, not to one word, so a stressed word needs its own sentence or a rewrite. Test two phrasings with a short script first: at $0.0475 per 1,000 characters, a 100-character test line is about half a cent.

Plan the split

Because the numeric ranges are fixed, you can check a script against them before any money moves. A speed of 1.1 is inside 0.6 to 1.5. A volume of 2.5 is outside 0.5 to 2.0 and the request will not be accepted. Keep the control values in your own config, not in the script text, so a markup-heavy script becomes plain text plus a settings object.

If a script mixes languages, set the language per request and split the text accordingly, because the field names the language the voice speaks the whole transcript in.

A last practical point: store the settings object with the job. The result of a TTS job can be read back later, so a team can see which speed and volume produced a take that sounded right and reuse them for the next episode.

One more difference to plan for: a voice model with markup support lets the writer shape a sentence in place. With typed fields, the shaping happens per request, so split a script into sentences when two parts of it need different speeds, and join the takes afterwards with Timeline audio, a flat $0.01 per job. Sume TTS segmentation can return the sentence timings so the joined file stays traceable.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume