Bulk text to speech from a CSV: one job per row, safe retries
Voice 500 CSV rows with Sume TTS 1.0. One idempotency key per row id, retry-after on 429, and the queue-capacity numbers per plan decide how fast it goes.

A spreadsheet of product lines, announcements or lesson sentences is the usual starting point for a bulk voiceover. The risk is not the TTS call. It is the retry: a script that crashes at row 312 and restarts from row 1 pays for 311 rows twice. Sume TTS 1.0 gives you what you need to avoid that, an Idempotency-Key header and durable jobs.
Use the row id as the key
Build the key from a stable value in the sheet, such as vo-2026-10- plus the row id. The jobs guide says a retry with the same key returns the original job instead of billing a second one. The same key with a different payload returns 409 idempotency_conflict. That is useful: if someone edited a row's text after the first submit, you learn about it instead of silently getting the old audio.
Know the queue before you launch
Per the generation admission page, jobs wait in a queue while a plan-limited number run at once. Submitting 500 rows is legal. Anything beyond the queue capacity is rejected with 429 queue_full, and 429 rate_limited carries a retry-after header you should honor.
| Plan | Running at once | Queue capacity |
|---|---|---|
| Free | 1 | 5 |
| Pro | 4 | 20 |
| Startup | 8 | 40 |
| Scale | 20 | 100 |
| Enterprise | 20 | 100 |
Code
Submit in waves no larger than your queue capacity, so queue_full never fires. This sketch uses the same submit helper as the other posts, without the polling loop, to stay short.
import csv, os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
rows = list(csv.DictReader(open("lines.csv", encoding="utf-8")))
def submit(row):
body = {"transcript": row["text"], "language": row["lang"], "mode": "async",
"voice": {"id": row["voice_id"]}, "metadata": {"row": row["id"]}}
while True:
r = requests.post(B + "/tts-1.0/generate", json=body,
headers={**H, "Idempotency-Key": "vo-2026-10-" + row["id"]})
if r.status_code == 429:
time.sleep(int(r.headers.get("retry-after", "5")))
continue
r.raise_for_status()
return r.json()["data"]["job"]["id"]
ids = {row["id"]: submit(row) for row in rows[:20]}
print(ids)Per-row failures
A bad row should not stop the batch. A voice and language pair that Sume flags returns 409 tts_voice_language_mismatch. A row longer than 20,000 characters is a 400. Write both to an errors column and move on. Rows whose audio is over 1,200 seconds fail the job with tts_duration_exceeded, with no credit captured.
What it costs
At $0.0475 per 1,000 characters, a 500-row sheet of 200-character lines is 100,000 characters, or $4.75 before per-job rounding. Each job quotes up to a whole cent, so 200 characters (about 0.95 cents) quotes at 1 cent, and 500 rows come to $5.00. Use the dry-run approach in the spend cap guide before the real run.
Sources
Related posts
More in Developers
- Bun script: submit a Wan 3.0 video job, poll it and save clip.mp4
A single Bun file that posts to /v1/videos for wan-3.0, polls until completed and writes clip.mp4 with Bun.write; the 5-second 480p test costs $0.3125.
- Burn an 'AI-generated' disclosure line into a clip with caption cues
YouTube asks creators to disclose realistic AI content. Add an authored on-screen line with Sume caption cues, no speech-to-text, that travels with the file.
- Burn captions from TTS word timings: no second transcription
Ask Sume TTS for word timings, group them into phrases and send them as caption cues. The caption job skips speech-to-text and still costs $0.20.
- Callback or polling for Omni 4K jobs on the Sume video API
Use callback_url on /v1/videos or mode webhook on the Video Router for Omni 4K batches: one signed request per job, not repeated polls. Checks included.
Written by Sume