Slow down a TTS voiceover: speed 0.6 to 1.5, volume, emotion

Sume TTS accepts generation_config with speed 0.6 to 1.5, volume 0.5 to 2 and a free-text emotion guide. How to retime a voiceover to a video.

5 min readSume
All posts

Sume TTS slows or speeds a voice through generation_config.speed, a multiplier from 0.6 to 1.5, and sets level with volume from 0.5 to 2. An emotion string of up to 64 characters guides the delivery. All three are optional and sit in one object, as the API reference schema shows.

Use speed to fit a script to a video length before you render, not to rescue a bad script afterwards. A cheap test render shows the effect, and each render is a new job.

How do the controls behave?

The schema lists volume as a multiplier in [0.5, 2.0], speed as a multiplier in [0.6, 1.5], and emotion as a free-text guide. It does not promise that two renders match sample for sample, and it does not list emotion values, so treat the emotion field as a hint and listen to the result.

generation_config fields (Sume OpenAPI schema, checked 2026-10-10)
FieldRangeNotes
speed0.6 to 1.5Multiplier; 1.0 is normal
volume0.5 to 2Multiplier
emotion1 to 64 charactersFree-text guide

Fit a script to a length

Take the duration of a first render, divide by the target, and adjust. A 40 second read that must fit 32 seconds needs a multiplier near 1.25, and one that must stretch from 24 to 30 seconds needs about 0.8. Both are in range. Beyond roughly 1.3 or below 0.7, rewrite the script instead; extreme speeds usually sound less natural.

Examples (arithmetic, not a quality claim)
Current readTargetSpeed multiplier
40 s32 s40 / 32 = 1.25
24 s30 s24 / 30 = 0.8
50 s60 s0.83
45 s30 s1.5 (the maximum)

What does it cost?

Speed and volume do not change the price, which is per input character: $0.0475 per 1,000, rounded up to a cent per job. So a 900-character script is $0.04275 before rounding, which becomes $0.05. Retiming by re-rendering therefore costs a few cents per attempt.

  • Test on one sentence first; it is the cheapest check.
  • Keep the speed value with the job record.
  • Use timings (timestamps.words) after you settle the speed, so captions match the final audio.

When to cut instead

If the voice is right but the take is too long by a few tenths of a second, trim with timeline audio. A change of speed alters pitch feel and pacing across the whole read; a trim only removes time. For pauses, use punctuation in the transcript and check the result.

A tested way to hit a target length

Do this in three renders. First, render at speed 1.0 and read the duration from the result. Second, compute the multiplier and render again. Third, if the second render still misses by more than a quarter second, trim with timeline audio instead of changing speed again. Each render of a 900-character script costs $0.05, so the whole test is about fifteen cents.

Keep the transcript constant while you test. If you edit words and change speed at the same time, you will not know which change fixed the length. Once the speed is chosen, request word timings for captions so the cues match the final audio.

  • Render 1: speed 1.0, read the duration.
  • Render 2: use the computed multiplier.
  • Render 3 or trim: only if the miss is large.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume