Bulk text to speech from a CSV: one job per row, safe retries

Voice 500 CSV rows with Sume TTS 1.0. One idempotency key per row id, retry-after on 429, and the queue-capacity numbers per plan decide how fast it goes.

5 min readSume
All posts

A spreadsheet of product lines, announcements or lesson sentences is the usual starting point for a bulk voiceover. The risk is not the TTS call. It is the retry: a script that crashes at row 312 and restarts from row 1 pays for 311 rows twice. Sume TTS 1.0 gives you what you need to avoid that, an Idempotency-Key header and durable jobs.

Use the row id as the key

Build the key from a stable value in the sheet, such as vo-2026-10- plus the row id. The jobs guide says a retry with the same key returns the original job instead of billing a second one. The same key with a different payload returns 409 idempotency_conflict. That is useful: if someone edited a row's text after the first submit, you learn about it instead of silently getting the old audio.

Know the queue before you launch

Per the generation admission page, jobs wait in a queue while a plan-limited number run at once. Submitting 500 rows is legal. Anything beyond the queue capacity is rejected with 429 queue_full, and 429 rate_limited carries a retry-after header you should honor.

Sume generation admission limits by plan (docs.sume.com, read 2026-10-05)
PlanRunning at onceQueue capacity
Free15
Pro420
Startup840
Scale20100
Enterprise20100

Code

Submit in waves no larger than your queue capacity, so queue_full never fires. This sketch uses the same submit helper as the other posts, without the polling loop, to stay short.

import csv, os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
rows = list(csv.DictReader(open("lines.csv", encoding="utf-8")))

def submit(row):
    body = {"transcript": row["text"], "language": row["lang"], "mode": "async",
            "voice": {"id": row["voice_id"]}, "metadata": {"row": row["id"]}}
    while True:
        r = requests.post(B + "/tts-1.0/generate", json=body,
                          headers={**H, "Idempotency-Key": "vo-2026-10-" + row["id"]})
        if r.status_code == 429:
            time.sleep(int(r.headers.get("retry-after", "5")))
            continue
        r.raise_for_status()
        return r.json()["data"]["job"]["id"]

ids = {row["id"]: submit(row) for row in rows[:20]}
print(ids)

Per-row failures

A bad row should not stop the batch. A voice and language pair that Sume flags returns 409 tts_voice_language_mismatch. A row longer than 20,000 characters is a 400. Write both to an errors column and move on. Rows whose audio is over 1,200 seconds fail the job with tts_duration_exceeded, with no credit captured.

What it costs

At $0.0475 per 1,000 characters, a 500-row sheet of 200-character lines is 100,000 characters, or $4.75 before per-job rounding. Each job quotes up to a whole cent, so 200 characters (about 0.95 cents) quotes at 1 cent, and 500 rows come to $5.00. Use the dry-run approach in the spend cap guide before the real run.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume