A 60,000-character narration on Sume TTS: 3 jobs, $2.85

Gradium tags 221 voices for narration. For long scripts on Sume TTS, here is the 20,000 character cap, the 1,200 second limit and the cost of 3, 4 or 6 jobs.

5 min readSume
All posts

A 60,000-character script is three Sume TTS jobs of 20,000 characters each and costs $2.85 at $0.0475 per 1,000 characters. Smaller splits cost a few cents more, because each job rounds up to a whole cent: four jobs of 15,000 characters are $2.88, and six jobs of 10,000 are also $2.88.

Gradium's October 7 page says 221 of its voices are tagged for narration but states no price, so there is nothing to compare it with here (read 2026-10-10 at Gradium). The Sume numbers come from the TTS price in its catalog and the request limits in the API reference.

What are the limits per job?

Sume TTS accepts up to 20,000 characters in transcript, and spaces and punctuation count. Separately, synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, and no credit is captured for that failure. The two caps are independent: a script of 20,000 characters read slowly, or with long pauses, can reach the duration cap first, so split by scene or chapter rather than at the exact character limit.

Send either transcript or transcript_source, never both.

What does each split cost?

The catalog states the rate as $0.0475 per 1,000 characters, and the billable formula as list times margin, rounded up to a whole cent per job. The table works the arithmetic for 60,000 characters.

Cost of a 60,000-character script by split (rate from the Sume catalog, checked 2026-10-10)
SplitCharacters per jobPer-job chargeTotal
3 jobs20,000$0.95$2.85
4 jobs15,000$0.7125 rounded up to $0.72$2.88
6 jobs10,000$0.475 rounded up to $0.48$2.88
12 jobs5,000$0.2375 rounded up to $0.24$2.88

How do you keep the voice consistent across jobs?

Use the same voice selector, language and engine id in every job. Pin the engine with the TTS Router (sonic-3.6, not sonic-latest) if you want the same model next month, because sonic-latest is an alias for the current stable Sonic release. Ask for wav if you will join the parts: timeline audio concatenates up to 20 parts sample by sample with no silence at the seams, and its docs recommend wav for anything you will join again, since mp3 adds priming padding at every edge.

  • Same voice, same language, same engine id in every part.
  • Split at paragraph or scene boundaries, not mid-sentence.
  • Concat with operation: "concat": $0.01 flat per job.

What does the full pipeline cost?

Three narration jobs ($2.85) plus one concat job ($0.01 flat) is $2.86 for one gapless wav, before any video, captions or music. If the concat is the 20-part maximum, the same $0.01 applies, so the join is not the place where cost grows. Check each rate in GET /v1/catalog before budgeting, since Sume's docs say to confirm live rates there.

What can go wrong on a long script?

Three failures are worth planning for. First, tts_duration_exceeded if one part synthesizes past 1,200 seconds; no credit is captured, so retry with a smaller part. Second, invalid_voice_id if the voice id has the wrong shape, which is rejected before any credits are reserved. Third, a language mismatch warning if your voice and language disagree; resend with confirm_language_mismatch only after a person has accepted it.

Keep a small table of part number, character count and job status in your own system. Because each job is independent, a failed part can be resubmitted alone. The jobs guide describes the status values to poll.

  • Retry only the failed part, with the same voice, language and engine id.
  • Do not change the engine id between parts of one narration.
  • Store the part order next to each result URL so the concat input order is explicit.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume