One narrator for 8 episodes: reuse the settings a Sume TTS job logs
Keep one voice across a series: send the same voice id, language and speed each episode, and read them back from the finished text_to_speech job. Costs $0.95.

To keep one narrator across a whole series on Sume, send the same voice.id, language and speed on every POST /v1/tts-1.0/generate call, and when in doubt read those values back from the finished episode-one job. A completed text_to_speech job records its voice, language, output format, generation_config and speed, and the docs say to read them from the job "to make the next line sound the same".
Eight episodes of 2,500 characters each is 20,000 characters of narration, which is $0.95 at the published $0.0475 per 1,000 characters (read 2026-10-03). Series are getting funding attention: Metricool reports TikTok's "The Next Episode" program with Amplify for creator-led series (reported, read 2026-10-03). A voice that drifts between episode 3 and episode 4 is the kind of flaw viewers notice first.
What the job record gives you
The request body for Sume TTS 1.0 takes transcript (up to 20,000 characters), a voice object with an id, an optional language, a speed enum of slow, normal or fast (the API marks it deprecated in favour of generation_config.speed, a number from 0.6 to 1.5), output_format, generation_config, and timestamps. Anything you do not send comes back as null on the finished job. That null is useful: it tells you which settings the engine chose for you, and a null on a field you assumed you had pinned is the bug.
Treat episode one as the reference. After it completes, read the record once, store it next to the series bible, and build every later request from that stored object instead of from memory. A narrator drifts when one episode is made by a script and another by hand with a different speed.
| Field | Values | Why pin it |
|---|---|---|
| voice.id | A voice id from your voices list | The identity of the narrator |
| language | A code such as en | Avoids a second language guess per episode |
| speed | slow, normal or fast | Pace is the first thing an audience hears change; the enum is deprecated, so prefer generation_config.speed |
| generation_config | Volume 0.5-2, speed 0.6-1.5, emotion; or null | Reuse what episode one actually used |
| output_format | Container, sample rate, bit rate | Default is mp3, 44100 Hz, 128000 bit/s; joins sound cleaner at one format |
Pin the engine too
Sume TTS 1.0 picks the engine for you. If you want the engine fixed for the whole season, use the TTS router, which requires a model taken from GET /v1/tts-router/models and then runs exactly that id. Voice selection, transcript rules and billing are the same as TTS 1.0; the engine picker is the only difference. Choose one id for the season and write it in the series bible.
The voices list is the place to find a voice id. If a request carries a voice name from another vendor you get a 400 invalid_voice_id, so copy the id, not the display name.
A runnable season loop
The script submits episode one, waits for the job, reads its record, and then reuses the voice, language and speed for episode two. It uses Idempotency-Key so a retried submit returns the original job instead of billing twice.
import os, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def tts(text, key, **cfg):
body = {"transcript": text, **cfg}
r = requests.post(f"{API}/v1/tts-1.0/generate",
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
return r.json()["request_id"]
def wait(job_id):
while True:
s = requests.get(f"{API}/v1/jobs/{job_id}/status", headers=H).json()
if s.get("status") == "completed":
res = requests.get(f"{API}/v1/jobs/{job_id}/result", headers=H).json()
return res.get("result", res)
if s.get("status") in ("failed", "canceled"):
raise RuntimeError(s)
time.sleep(3)
ep1 = tts("Last night the lights went out on Mercer Street.", "s1e1-vo",
voice={"id": os.environ["NARRATOR_VOICE_ID"]}, language="en", speed="normal")
rec = wait(ep1)
keep = {k: rec[k] for k in ("voice", "language", "speed", "generation_config") if rec.get(k)}
print("reusing", keep)
ep2 = tts("By morning, nobody on the street would speak of it.", "s1e2-vo", **keep)
print("episode 2 job", ep2)What to do before the season starts
Make episode one's narration first and listen to it before you queue the rest. If you are unhappy with the pace, change speed now, because changing it on episode five means re-recording episodes one to four.
Cut the script at paragraph boundaries and keep each request well under the 20,000 character cap, so a failed line costs a few cents instead of a full episode. Join the lines with the audio parts feature in Timeline rather than re-speaking them.
Name every request after its episode, as in s1e2-vo above. With an Idempotency-Key per episode, a network retry returns the original job instead of creating and billing a second one, and a re-record of a single episode is a deliberate new key such as s1e2-vo-take2. Keep the keys in your episode table so a rerun of the whole season script is safe. Last, add up the bill before you press go: 8 requests of 2,500 characters is 8 x $0.11875, which is $0.95, so narration is not where a season budget goes.
What Sume does and does not do
Sume lets you fix the voice id, language, speed and, through the router, the engine, and it records what was used. It does not guarantee that a voice sounds identical across two different engines, so do not switch engines mid-season.
Sources
Related posts
More in Use cases
- Saved favorites to a lookbook sheet: one gpt-image-2.5 call
Turn up to 16 saved product images into one lookbook sheet with openai/gpt-image-2.5 on Sume. Reference numbering, layout wording and what to check.
- School announcement video in two languages for parents: captions
One clean announcement clip, two caption jobs: English and your second language. Language hints, Korean styles and a human check on the translation.
- Screen-recording tutorial Shorts: what YouTube expects you to add
A screen recording with no voice is thin under YouTube's originality rules. How to cut a tutorial Short from your own recording with trim and caption cues.
- Section 508 caption display rules: two lines, 45 characters, burned in
Section508.gov asks for two lines, 45 characters per line, white text on a translucent black box. How each rule maps to Sume's caption design fields.
Written by Sume