Slow narration for language learners: what speed 0.6 does to cost
Slow TTS narration to speed 0.6 for learners. The price stays per character, but the audio gets longer, which moves the render and caption steps after it.

Slowing narration for language learners costs nothing extra at the speech step, because Sume bills text to speech per transcript character, not per second: a 600-character lesson is about $0.0285 at speed 0.6 and at speed 1.0 alike. What changes is the length. generation_config.speed runs from 0.6 to 1.5, and a slower read makes a longer file, so the steps that price by length, the Timeline render and the $0.20 caption job, can move. Send language and a voice recorded in it, request timestamps.words, and read the real length from the job before you plan the rest.
The speed range, the word timings and the price per character come from the Sume API reference and API pricing; the render, caption and polling rules from Timeline 1.0, Video captions and Jobs and results, all read on 2026-10-03. The per-sentence clips for drills are in Text to speech for language learning; this page is about what slow speed does to length and cost.
Does a slower voice cost more?
Not on the speech step. The pricing page rates Sume TTS 1.0 at $0.0475 per 1,000 characters, spaces and punctuation included, with a 20,000-character request maximum. Speed is not a billing input. A slow take and a normal take of the same passage are two requests, each billed on its own characters, so only a retake costs more. Respelling words for clarity changes the character count, so count after you edit.
What does 0.6 do to the length?
It lengthens the audio. Sume's docs state the multiplier range but not an exact ratio between speed and seconds, so do not divide by 0.6 and trust it. With timestamps.words: true the completed result carries duration_seconds and words[] with each word's start and end, so submit once and read the real length. The word timings come back in seconds on the slowed audio, ready for captions.
Which later steps depend on that length?
A learner clip is usually the narration over a still, then captions. Both steps price by length, so a slower read can change the bill even though the voice did not.
- A 55-second read that slows to 65 seconds renders as two minutes, $0.20 instead of $0.10.
- The caption price is documented only up to 60 seconds. Past that, split the clip or see Add captions to a long video by API.
| Step | Priced by | Moves with speed? |
|---|---|---|
| Text to speech | $0.0475 per 1,000 characters | No |
| Timeline render | $0.10 per output minute, rounded up | Yes, through length |
| Video captions | $0.20 for videos up to 60 seconds | Yes, through length |
| Word timings | Seconds on the slowed audio | Yes, they stretch with it |
How do I request slow narration and read its length?
Set voice.id to a voice recorded in the lesson's language, otherwise the 409 tts_voice_language_mismatch check can stop the submit. This script submits one slow take, polls it, and prints the cost and the measured length. Set SUME_API_KEY and SUME_VOICE_ID first.
import os, time, requests
API, H = "https://api.sume.com", {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
text = "¿Dónde está la estación? Está a dos cuadras."
body = {"transcript": text, "language": "es", "voice": {"id": os.environ["SUME_VOICE_ID"]},
"generation_config": {"speed": 0.6}, "timestamps": {"words": True},
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}}
r = requests.post(f"{API}/v1/tts-1.0/generate", json=body,
headers={**H, "Idempotency-Key": "lesson-07-slow"})
r.raise_for_status()
d = r.json()["data"]
while not requests.get(d["status_url"], headers=H).json()["data"]["terminal"]:
time.sleep(2)
res = requests.get(d["result_url"], headers=H).json()["data"]["result"]
print(len(text), "characters =", round(len(text) * 0.0475 / 1000, 5), "USD")
print("audio", res["duration_seconds"], "seconds,", len(res["words"]), "words")How do I burn the captions afterwards?
Caption jobs take a video, so lay the narration over a still with Timeline first, then send the word timings as words with text, start and end. The TTS result names each word word, so rename the key when you map it. Burn the captions after the render, from the audio you already measured, and nothing is transcribed twice.
Sources
Related posts
- Text to speech for language learning: slow sentence audio
- Add captions to a long video by API: split, caption, rejoin
- TTS speed: slow, normal, fast or generation_config.speed 0.6 to 1.5?
- TTS word timings to burned-in captions: send them as words on Sume
- Text to speech time calculator: estimate, then measure
- How to burn captions onto a video with the Sume API
More in Use cases
- Small Business Saturday 2026 video: shelf photos to a 9:16 ad
Small Business Saturday is Nov 28, 2026. Turn five shelf photos into five clips, add a date card, and join them in one render, with limits and a plan.
- Small Business Saturday for service shops: one offer clip per type
Small Business Saturday is November 28, 2026, and the SBA says to tailor offers by business type. Four service-shop offer clips cost $0.88 on Sume.
- Snapchat AI Clips make 5-second videos; do it outside Snapchat
Snap's AI Clips turn a photo into a five-second video inside a Lens. Here is how to get the same photo-to-clip result as a file you can post anywhere.
- Snapchat Commercials: the first 6 seconds can't be skipped
Snapchat lists Commercials as 9:16, 720x1280, H.264 MP4 or MOV, 3 to 180 s, first 6 s non-skippable. Put the brand and offer in those 6 seconds.
Written by Sume