Slow breathing voiceover: Sume TTS speed 0.6 is the floor

For a calm, slow voiceover set generation_config.speed as low as 0.6 on Sume TTS. Below that the request is rejected, so slow the rest through writing.

5 min readSume
All posts

To slow a Sume TTS voice for a breathing or meditation clip, set generation_config.speed to a value between 0.6 and 1.5. 0.6 is the slowest the API accepts. A value outside the range fails validation instead of being clamped, so a slower read than 0.6 has to come from the script, not from the speed field.

The ranges are in the TTS 1.0 request schema in the Sume API reference (read 2026-10-06).

What does the speed field accept?

generation_config holds three optional controls. They apply to the voice, so they work with an avatar voice (avatar_handle) or a voice.id.

TTS 1.0 generation_config ranges (read 2026-10-06)
FieldRangeMeaning
speed0.6 to 1.5Speed multiplier; 1.0 is normal
volume0.5 to 2.0Volume multiplier
emotion1 to 64 charactersOptional emotion guide

How do I get slower than 0.6?

You cannot with this field, and the schema has no SSML or pause tag. The levers left are in the text: short sentences, one idea each, and the sentence segmentation option. With timestamps.words set to true and segmentation.mode set to sentence, the result returns gapless per-sentence segments, and with a WAV output each segment carries its own audio file. You can then place each sentence on your own timeline with the space you want between them.

What should I check before I ship it?

Read words[] from the result and compare the total length with the video. At 0.6 a script reads much longer than at 1.0, and a TTS job fails with tts_duration_exceeded if the synthesized audio is longer than 1,200 seconds, with no credit captured. Listen to the first take at 0.6, 0.8 and 1.0 before you queue the whole series.

Sources

More in Media tools

All Media tools posts

Written by Sume