Slow, clear learner audio: Sume TTS speed 0.6 and sentence slices

Sume TTS speed goes down to 0.6. For learner audio, render at 0.8, slice by sentence with word timings, and keep the full pass for comparison. Cost per lesson.

4 min readSume
All posts

Learners want audio that is slower than native speech but not robotic. Sume TTS exposes generation_config.speed from 0.6 to 1.5, so you can render slower audio without editing a file.

Pick the speed

Suggested speeds for learner audio, my rules of thumb, read 2026-10-05
Learner levelSpeedWhy
Absolute beginner0.7Each syllable is audible
Beginner0.8Slower, rhythm intact
Intermediate0.9Near natural
Shadowing practice1.0Match real speech

Slice per sentence

Ask for segmentation: {mode: "sentence"} with timestamps.words true and a wav container. You get one clip per sentence and the timings, so the app can play a single sentence on tap. mp3 returns timings only, no per-slice audio URLs. Lines are gapless, so add a short pause in the player if you want repeat-after-me gaps.

Set the language

Always send language for a non-English lesson, and a voice tagged for that language. Sume answers a known mismatch with a 409 (tts_voice_language_mismatch) before any job or charge, and you retry with confirm_language_mismatch: true only if you meant it. Speed 0.6 on a wrong-language voice sounds worse, not clearer.

Cost

A 600-character lesson is $0.0285 raw, so 3 cents billed. Render it at 0.8 and 1.0 and the pair is 6 cents. A 30-lesson course with both versions is under $2.

Related posts

More in Use cases

All Use cases posts

Written by Sume