100 holiday voice-overs at 600 characters: $3.00 of Sume TTS
Sume TTS bills $0.0475 per 1,000 characters, rounded up per job. One 600-character read is $0.03, so 100 holiday ad voice-overs cost $3.00. Script included.

A hundred 600-character holiday ad voice-overs cost $3.00 on Sume TTS. The rate is $0.0475 per 1,000 characters, so one read is 600 × $0.0475 / 1,000 = $0.0285, which is billed as $0.03 after rounding up to the cent.
Spaces and punctuation count as characters, and a single transcript can be up to 20,000 characters. That makes a voice-over the cheapest step in a holiday ad pipeline, and rounding is the only thing worth watching.
How rounding changes the total
Without rounding, 100 reads of 600 characters are 60,000 characters, or 60 × $0.0475 = $2.85. With per-job rounding each read is $0.03, so the batch is $3.00. The extra $0.15 is the rounding. If you merge six short lines into one job of 600 characters, you pay $0.03 for one job, not 6 × $0.01 = $0.06 for six.
| Script length | Raw cost | Billed per job | 100 jobs |
|---|---|---|---|
| 300 characters | $0.0143 | $0.02 | $2.00 |
| 600 characters | $0.0285 | $0.03 | $3.00 |
| 1,000 characters | $0.0475 | $0.05 | $5.00 |
| 2,700 characters (3 minutes) | $0.1283 | $0.13 | $13.00 |
| 20,000 characters (maximum) | $0.95 | $0.95 | $95.00 |
Count before you call
Put the counting in code so a budget check runs before the first paid call. This script totals the batch with the same rule as above.
import math
RATE_PER_1000 = 0.0475
def job_cost(text: str) -> float:
raw = len(text) * RATE_PER_1000 / 1000
return math.ceil(round(raw * 100, 6)) / 100
scripts = ["Last order day for gifts is the 18th. " * 15] * 100
lens = {len(s) for s in scripts}
total = sum(job_cost(s) for s in scripts)
print(f"{len(scripts)} jobs, lengths {sorted(lens)}, total ${total:.2f}")
Mistakes that cost more than the characters
The voice must match the language. If the language you send differs from the primary language of the voice, TTS returns 409 tts_voice_language_mismatch, and you pick another voice rather than retrying. A script that runs longer than the video it goes under forces a re-render of the clip, which costs far more than the voice.
Use an Idempotency-Key for each read, and take the result by webhook when you run a hundred at once. The jobs and webhooks docs describe the status route and the signed terminal event.
Where TTS sits in the whole ad
At $0.03 a read, voice is about 4% of a $0.75 six-second Omni Flash clip at 720p. The big lines are video seconds and avatar seconds, so spend your attention there, and use TTS freely to test script variants.
Sources
Related posts
More in Pricing
- 10,000 eight-second clips: STT billed per second vs per hour
Ten thousand 8-second clips are 80,000 seconds, or 22.2 hours: $12.00 on MAI-Transcribe-2-Streaming, $13.33 on Sume STT, which bills each job by the second.
- 10,000 images a month on Sume: GPT Image 2.5, Grok, Nano Banana 2
A month of 10,000 images costs $73.75 on GPT Image 2.5 low, $165 on medium, $250 on Grok Imagine and $1,000 on Nano Banana 2 1K on Sume. Table and script.
- 160 ads a month at 450 characters: the TTS bill on Sume, MAI and Flash
A month of 160 ad voiceovers at 450 characters costs $4.80 on Sume, $1.584 on MAI-Voice-2.1 and $1.08 on Flash. The arithmetic and what the gap buys.
- 20 ad hook variants of 5 seconds: cheapest Sume video model per batch
Twenty 5-second hook tests cost $3.80 on Omni at 360p, $6.40 on Wan at 480p and $12.60 at 720p on Omni or Wan. Per-job rounding is included.
Written by Sume