Sume TTS has no SSML field: speed, emotion, pronunciation dictionary
Moving an SSML voice script to Sume TTS? The request takes plain transcript text, so use speed 0.6 to 1.5, volume, emotion and a pronunciation dictionary id.

No. Sume's TTS request has no SSML or markup field. It takes plain transcript text (1 to 20,000 characters) and a small set of typed controls: generation_config.speed from 0.6 to 1.5, generation_config.volume from 0.5 to 2.0, a free-text emotion guide, a pronunciation_dict_id, and language. If your script is SSML, strip the tags and move each intent to one of those fields.
New voice models launched on 2026-10-01 at $22 per 1M characters (MAI-Voice-2.1) and $15 per 1M characters (Flash) (Microsoft AI, read 2026-10-05), so scripts written for them will come with markup on top of the words.
Where each intent goes
Each markup intent has a Sume field, a coarser one, or none. The table says which. It does not claim that the voices behave the same.
| Intent in a markup script | Sume TTS field | Range or form |
|---|---|---|
| Faster or slower delivery | generation_config.speed | 0.6 to 1.5 |
| Louder or quieter | generation_config.volume | 0.5 to 2.0 |
| Mood of the read | generation_config.emotion | Free-text guide |
| Name or term pronunciation | pronunciation_dict_id | One dictionary id |
| Language of the voice | language | BCP-47 or ISO-639, such as ko or ja |
| Exact pause length | None documented | Not a field |
Strip tags before you pay
Characters count whatever you send: the schema says spaces and punctuation count toward usage. Tags left inside the transcript would be billed as text and might be spoken, so remove them before you submit. On the MCP tool, dry_run previews the cost and max_spend_usd caps it.
Language and the older speed field
Set language for every non-English transcript. When it is omitted the provider defaults to English, with a fallback guess for Hangul-only or kana-only text. If the voice and language disagree, Sume returns tts_voice_language_warning, and you retry with confirm_language_mismatch: true only after the user agrees.
The old speed enum with slow, normal and fast is deprecated. Prefer the numeric field.
A request body
A minimal request body, with the controls in one place:
{
"idempotency_key": "ad-read-001",
"payload": {
"transcript": "Spring sale starts Friday. Everything is half price.",
"language": "en",
"voice": {"id": "voi_00000000000000000000000000000000"},
"generation_config": {"speed": 1.1, "volume": 1.0, "emotion": "warm"},
"output_format": {"container": "mp3", "sample_rate": 44100, "bit_rate": 128000}
}
}Notes
That is the MCP shape, with fields inside payload. The HTTP route takes the same fields at the top level. The voice id above is a placeholder, so copy a real id verbatim from your voice library or resolve one from an avatar handle.
Pauses are the gap. If a script depends on exact break lengths, write the pause into the text with punctuation and listen, because the schema offers no timing field for it.
The same goes for emphasis on a single word. The documented controls apply to the whole request, not to one word, so a stressed word needs its own sentence or a rewrite. Test two phrasings with a short script first: at $0.0475 per 1,000 characters, a 100-character test line is about half a cent.
Plan the split
Because the numeric ranges are fixed, you can check a script against them before any money moves. A speed of 1.1 is inside 0.6 to 1.5. A volume of 2.5 is outside 0.5 to 2.0 and the request will not be accepted. Keep the control values in your own config, not in the script text, so a markup-heavy script becomes plain text plus a settings object.
If a script mixes languages, set the language per request and split the text accordingly, because the field names the language the voice speaks the whole transcript in.
A last practical point: store the settings object with the job. The result of a TTS job can be read back later, so a team can see which speed and volume produced a take that sounded right and reuse them for the next episode.
One more difference to plan for: a voice model with markup support lets the writer shape a sentence in place. With typed fields, the shaping happens per request, so split a script into sentences when two parts of it need different speeds, and join the takes afterwards with Timeline audio, a flat $0.01 per job. Sume TTS segmentation can return the sentence timings so the joined file stays traceable.
Sources
Related posts
More in Developers
- Sume /v1/videos says cancelled, /v1/jobs says canceled: guard it
The /v1/videos poll uses pending, in_progress and cancelled; /v1/jobs uses queued, processing and canceled. A small normalizer keeps your poller from hanging.
- Sume webhook and status poll race: never move a job row backwards
A late poll can say processing after the webhook already said completed. A rank-guarded SQLite update keeps a Sume job row from going backwards.
- Sume webhook five-minute replay window: reject stale deliveries
Sume signs timestamp.raw_body and advises a five-minute replay tolerance. Reject stale timestamps, refuse an empty secret, and keep job_id as the dedupe key.
- Sume webhook retries: 10 attempts 30 seconds apart, size your downtime
Sume retries a failing webhook up to 10 attempts at a default 30-second spacing, about 4.5 minutes. Past that, redeliver or poll after a publisher outage.
Written by Sume