Slow narration for language learners: what speed 0.6 does to cost

Slow TTS narration to speed 0.6 for learners. The price stays per character, but the audio gets longer, which moves the render and caption steps after it.

4 min readSume
All posts

Slowing narration for language learners costs nothing extra at the speech step, because Sume bills text to speech per transcript character, not per second: a 600-character lesson is about $0.0285 at speed 0.6 and at speed 1.0 alike. What changes is the length. generation_config.speed runs from 0.6 to 1.5, and a slower read makes a longer file, so the steps that price by length, the Timeline render and the $0.20 caption job, can move. Send language and a voice recorded in it, request timestamps.words, and read the real length from the job before you plan the rest.

The speed range, the word timings and the price per character come from the Sume API reference and API pricing; the render, caption and polling rules from Timeline 1.0, Video captions and Jobs and results, all read on 2026-10-03. The per-sentence clips for drills are in Text to speech for language learning; this page is about what slow speed does to length and cost.

Does a slower voice cost more?

Not on the speech step. The pricing page rates Sume TTS 1.0 at $0.0475 per 1,000 characters, spaces and punctuation included, with a 20,000-character request maximum. Speed is not a billing input. A slow take and a normal take of the same passage are two requests, each billed on its own characters, so only a retake costs more. Respelling words for clarity changes the character count, so count after you edit.

What does 0.6 do to the length?

It lengthens the audio. Sume's docs state the multiplier range but not an exact ratio between speed and seconds, so do not divide by 0.6 and trust it. With timestamps.words: true the completed result carries duration_seconds and words[] with each word's start and end, so submit once and read the real length. The word timings come back in seconds on the slowed audio, ready for captions.

Which later steps depend on that length?

A learner clip is usually the narration over a still, then captions. Both steps price by length, so a slower read can change the bill even though the voice did not.

  • A 55-second read that slows to 65 seconds renders as two minutes, $0.20 instead of $0.10.
  • The caption price is documented only up to 60 seconds. Past that, split the clip or see Add captions to a long video by API.
What speed changes downstream, from Sume's pricing page and docs, read 2026-10-03.
StepPriced byMoves with speed?
Text to speech$0.0475 per 1,000 charactersNo
Timeline render$0.10 per output minute, rounded upYes, through length
Video captions$0.20 for videos up to 60 secondsYes, through length
Word timingsSeconds on the slowed audioYes, they stretch with it

How do I request slow narration and read its length?

Set voice.id to a voice recorded in the lesson's language, otherwise the 409 tts_voice_language_mismatch check can stop the submit. This script submits one slow take, polls it, and prints the cost and the measured length. Set SUME_API_KEY and SUME_VOICE_ID first.

import os, time, requests
API, H = "https://api.sume.com", {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
text = "¿Dónde está la estación? Está a dos cuadras."
body = {"transcript": text, "language": "es", "voice": {"id": os.environ["SUME_VOICE_ID"]},
        "generation_config": {"speed": 0.6}, "timestamps": {"words": True},
        "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100}}
r = requests.post(f"{API}/v1/tts-1.0/generate", json=body,
                  headers={**H, "Idempotency-Key": "lesson-07-slow"})
r.raise_for_status()
d = r.json()["data"]
while not requests.get(d["status_url"], headers=H).json()["data"]["terminal"]:
    time.sleep(2)
res = requests.get(d["result_url"], headers=H).json()["data"]["result"]
print(len(text), "characters =", round(len(text) * 0.0475 / 1000, 5), "USD")
print("audio", res["duration_seconds"], "seconds,", len(res["words"]), "words")

How do I burn the captions afterwards?

Caption jobs take a video, so lay the narration over a still with Timeline first, then send the word timings as words with text, start and end. The TTS result names each word word, so rename the key when you map it. Burn the captions after the render, from the audio you already measured, and nothing is transcribed twice.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume