Language listening practice: one passage at 0.6, 1.0 and 1.5 speed
Make slow, normal and fast versions of one passage with Sume TTS generation_config.speed (0.6 to 1.5). Three jobs for 9 cents on a 600-character text.

Run the same passage through Sume TTS three times with generation_config.speed set to 0.6, 1.0 and 1.5, and you have slow, normal and fast listening files. For a 600-character passage each job is 3 cents, so the set is 9 cents.
The speed range in the API reference is 0.6 to 1.5. Speed changes how long the audio runs, not how many characters are billed.
The settings that matter
The older speed enum of slow, normal and fast is deprecated. Use the numeric field inside generation_config.
| Field | Range or form | Use in a lesson |
|---|---|---|
| generation_config.speed | 0.6 to 1.5 | Slow for first listen, 1.0 for review, 1.5 for stretch |
| generation_config.volume | 0.5 to 2.0 | Even out a quiet voice |
| generation_config.emotion | Free string, up to 64 characters | Calm, warm or curious delivery |
| language | BCP-47 code | Match the voice to the text |
| timestamps.words | On or off | Word timings for highlighting |
A workable recipe
Keep the text and voice identical across the three jobs so the only difference is speed.
- Pick a voice that supports your target language. A mismatch returns 409
tts_voice_language_mismatch. - Set
languageto the BCP-47 code of the passage, for example one for the language being taught. - Send three requests with speeds 0.6, 1.0 and 1.5, each with its own
Idempotency-Keyso a retry does not double-bill. - Request word timestamps on the slow version, so a learner app can highlight words as they play.
- Name the files by speed, and keep the transcript with them.
Cost for a course
At $0.0475 per 1,000 characters and a round-up per job, a 600-character passage is 3 cents per speed. Twenty passages at three speeds is 60 jobs and $1.80. Short phrases of under 210 characters hit the 1-cent minimum, so group sentences into passages instead of making a file per phrase.
If the lessons are in several languages, make sure each text is under the 20,000-character cap and that audio stays under the 1,200-second limit, which a lesson passage will not approach.
What to say to learners
A synthetic voice is a practice aid, not a native speaker. Label the audio as generated in the lesson, and let a human check the first batch of any language you cannot read yourself. Sume does not claim the voices match any regional accent, so listen before you publish.
Sources
Related posts
More in Use cases
- LinkedIn 'Seems like AI slop' button: what AI video posters change
LinkedIn lets members flag posts as AI slop. What that means for AI video on the feed, and a short Sume checklist to keep the work specific and yours.
- LinkedIn sorts posts spam, low-quality or clear: batching AI video
LinkedIn says classifiers label each post spam, low-quality or clear in near real time. How to vary a batch of AI videos so each one stands alone.
- Lip sync AI music video: split a song into 5 to 14.8 second windows
MiniMax H3 Max lip sync accepts 5 to 14.8 seconds of audio per clip. Split a 60-second song into five 12-second ranges: about $6.24 at 768p on Sume.
- Listing walkthrough from 8 photos: seven 4-second Wan 3.0 transitions
Eight room photos become seven Wan 3.0 first-to-last-frame transitions of 4 seconds. $3.50 at 720p plus $0.10 for the timeline render on Sume.
Written by Sume