SSML in text to speech: Sume takes a plain transcript, no ssml field

Does Sume's text to speech accept SSML? The tts_create body has a plain transcript and rejects unknown keys. What to use for speed, volume, emotion and pauses.

4 min readSume
All posts

Sume's hosted tts_create tool has no SSML field. Its body takes a plain transcript of up to 20,000 characters, and a key outside the documented set fails with unsupported_payload_keys instead of being passed through. Speed, volume and emotion are set with a generation_config object, and word and sentence timings come from timestamps and segmentation. Fields come from the tool's payload schema, summarized in MCP tools and gates.

What replaces the SSML tags

Most people reach for SSML for three things: pace, loudness and pauses. Sume covers the first two with numbers and leaves pauses to the text.

SSML habits and the Sume field that covers them, per the tts_create schema read 2026-10-03
You wantedSume fieldRange or form
prosody rategeneration_config.speed0.6 to 1.5
prosody volumegeneration_config.volume0.5 to 2
Emotiongeneration_config.emotionFree text, 1 to 64 characters
Sentence boundariessegmentation.modesentence, with boundary_lead_ms 0 to 500
Word timingtimestamps.wordstrue returns words[] with start and end seconds
Break or pausePunctuation in the transcriptNo pause tag in the documented body

Sending a request that is accepted

Keep the transcript to the words you want spoken, and do not send markup the schema does not list. If you migrated a script from another engine, strip its markup first and re-create the pace with speed.

The deprecated speed string (slow, normal, fast) still exists, but the schema says to prefer generation_config.speed. Once the job completes, the job result echoes the settings it used, so you can read them back, as the Jobs and results guide describes.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume