A 1,850-character narration in Sume TTS: price, and what 3 takes add

A 1,850-character script costs $0.0879 in Sume TTS at $0.0475 per 1,000 characters; three takes cost $0.2636. Table by length, and how characters count.

4 min readSume
All posts

Sume bills text-to-speech at $0.0475 per 1,000 characters, so a 1,850-character narration costs $0.0879 and three takes of it cost $0.2636. The unit is characters of the transcript, not seconds of audio.

The API reference says spaces and punctuation count toward usage, and one request takes a maximum of 20,000 characters. Count the string you send, not the words you typed.

Cost by script length

Per take = characters / 1,000 x $0.0475, from the price tool. Three takes is 3 x the single take.

Sume TTS price by transcript length (read 2026-10-09)
CharactersTypical useOne takeThree takes
600short caption read$0.0285$0.0855
185060-second narration$0.0879$0.2636
3700two-minute voiceover$0.1758$0.5272
7400four-minute explainer$0.3515$1.0545

Count before you send

A character counter in your pipeline gives the bill before the request exists. This works for any script and keeps retries honest, because a re-read of the same text costs the same again.

script = open("narration.txt", encoding="utf-8").read().strip()
chars = len(script)  # spaces and punctuation count
per_take = chars / 1000 * 0.0475
print(f"{chars} characters -> ${per_take:.4f} per take, ${3 * per_take:.4f} for 3")
assert chars <= 20000, "split the script: max 20,000 characters per request"

Where TTS sits in a video budget

A 1,850-character read is roughly a minute of speech, depending on the voice and pace. Against a 10-second Wan 3.0 720p clip at $1.25, the narration for a whole minute ($0.0879) costs about 7 percent of one clip. Voiceover is rarely the line to optimize; take count and clip retries are.

Splitting the script does not change the price

Billing is per character, so cutting a 3,700-character script into two 1,850-character requests costs $0.1758, the same as one request ($0.1758). Split on sentence boundaries when you want to re-read only the weak half: you pay for the half you redo, not the whole script.

That makes per-scene synthesis the cheap way to iterate. A producer who rejects scene 3 of 6 pays for scene 3 again, not for all six scenes.

Gotchas

  • Edit the script before you synthesize: each re-read bills the full character count again.
  • Characters are billed as sent, so remove stage directions and markdown markers you do not want spoken.
  • The TTS Router requires a model from GET /v1/tts-router/models; read the catalog rather than hard-coding an id.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume