Language listening practice: one passage at 0.6, 1.0 and 1.5 speed

Make slow, normal and fast versions of one passage with Sume TTS generation_config.speed (0.6 to 1.5). Three jobs for 9 cents on a 600-character text.

4 min readSume
All posts

Run the same passage through Sume TTS three times with generation_config.speed set to 0.6, 1.0 and 1.5, and you have slow, normal and fast listening files. For a 600-character passage each job is 3 cents, so the set is 9 cents.

The speed range in the API reference is 0.6 to 1.5. Speed changes how long the audio runs, not how many characters are billed.

The settings that matter

The older speed enum of slow, normal and fast is deprecated. Use the numeric field inside generation_config.

Sume TTS 1.0 listening-practice settings, from the API reference (read 2026-10-07)
FieldRange or formUse in a lesson
generation_config.speed0.6 to 1.5Slow for first listen, 1.0 for review, 1.5 for stretch
generation_config.volume0.5 to 2.0Even out a quiet voice
generation_config.emotionFree string, up to 64 charactersCalm, warm or curious delivery
languageBCP-47 codeMatch the voice to the text
timestamps.wordsOn or offWord timings for highlighting

A workable recipe

Keep the text and voice identical across the three jobs so the only difference is speed.

  • Pick a voice that supports your target language. A mismatch returns 409 tts_voice_language_mismatch.
  • Set language to the BCP-47 code of the passage, for example one for the language being taught.
  • Send three requests with speeds 0.6, 1.0 and 1.5, each with its own Idempotency-Key so a retry does not double-bill.
  • Request word timestamps on the slow version, so a learner app can highlight words as they play.
  • Name the files by speed, and keep the transcript with them.

Cost for a course

At $0.0475 per 1,000 characters and a round-up per job, a 600-character passage is 3 cents per speed. Twenty passages at three speeds is 60 jobs and $1.80. Short phrases of under 210 characters hit the 1-cent minimum, so group sentences into passages instead of making a file per phrase.

If the lessons are in several languages, make sure each text is under the 20,000-character cap and that audio stays under the 1,200-second limit, which a lesson passage will not approach.

What to say to learners

A synthetic voice is a practice aid, not a native speaker. Label the audio as generated in the lesson, and let a human check the first batch of any language you cannot read yourself. Sume does not claim the voices match any regional accent, so listen before you publish.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume