How much does a minute of AI voiceover cost? Measure it from one job
Cost per minute of voiceover depends on how fast the voice reads. Time one Sume TTS job, get characters per minute from word timestamps, and project the price.

Sume bills TTS by character, not by minute, so the price of a minute of voiceover depends on how many characters a minute of speech holds. Measure it instead of guessing: render one representative script, take the last word's end time from the word timestamps, and divide. The script below prints characters per minute and the cost per minute at $0.0475 per 1,000 characters (pricing, read 2026-10-06). Your own voice, speed and script decide the number, which is why the script measures it.
The comparison point is Microsoft's new pricing: MAI-Voice-2.1 at $22 per 1M characters and Flash at $15 per 1M (October tracker, read 2026-10-06). Per-character pricing makes minutes easy to convert once you have your own number.
Measure characters per minute
import json, os, time, urllib.request as u
KEY = os.environ["SUME_API_KEY"]
def call(url, body=None, key=None):
h = {"Authorization": "Bearer " + KEY, "Content-Type": "application/json"}
if key: h["Idempotency-Key"] = key
data = json.dumps(body).encode() if body else None
return json.load(u.urlopen(u.Request(url, data=data, headers=h)))["data"]
PRICE_PER_1K = 0.0475
SCRIPT = ("Our new grinder is ready in ten seconds. It fits in a drawer, runs quietly, "
"and cleans with one rinse. Order this week and get free shipping on every pack.")
job = call("https://api.sume.com/v1/tts-router/generate", {
"model": "sonic-3.6", "language": "en", "avatar_handle": os.environ["SUME_AVATAR_HANDLE"],
"transcript": SCRIPT, "timestamps": {"words": True}}, "cpm-measure-v1")
while not job["terminal"]:
time.sleep(job.get("next_poll_after_seconds") or 2)
job = call(job["status_url"])
words = call(job["result_url"])["result"]["words"]
seconds = max(w["end"] for w in words)
chars = len(SCRIPT)
cpm = chars / seconds * 60
print(f"{chars} chars in {seconds:.1f} s = {cpm:.0f} chars/min")
print(f"one minute costs ${cpm / 1000 * PRICE_PER_1K:.4f}")
print(f"a 10M char month costs ${10_000 * PRICE_PER_1K:,.2f}")Reading the result
Sume's catalog says spaces and punctuation count toward the billed characters, so the script uses len(SCRIPT). Non-Latin scripts and emoji can count differently from a Python length, so treat the result as an estimate and compare it with the charge on your first real job. Voice, language and speed all change the result: a slower speed setting means fewer characters a minute, which lowers the cost of a minute but not the cost of a script. Budget from characters, since the characters are what you are billed for, and use minutes only as a way of explaining the figure to someone else. Re-measure when you change voice or model. For a fixed script, the price is the same at any speed.
| Volume | Sume at $0.0475 per 1k | MAI at $22 per 1M | Flash at $15 per 1M |
|---|---|---|---|
| 1M characters | $47.50 | $22.00 | $15.00 |
| 10M characters | $475.00 | $220.00 | $150.00 |
| 100,000 characters | $4.75 | $2.20 | $1.50 |
Use the measured rate
The measured number changes the answer in a useful way. If your script reads slowly, you pay for fewer characters per minute of finished audio; if you fill every second with dense copy, you pay for more. A tighter script is a lower bill on any provider with per-character pricing, Microsoft included, and it makes shorter, better ads as well. Run the script on two or three representative scripts, not one, and use the average for the budget.
Where Sume is not the cheapest
Be plain about what the table says: Sume's per-character rate is higher than Microsoft's listed rates, and the tracker's figures are as read on 2026-10-06 and can change. The difference is what you get around the voice: one API key and one job pattern for TTS, transcription, captions, joins and video, with webhooks and idempotency built in. For most teams the voice line is a small part of a video budget. To estimate a script before you submit it, see TTS cost estimate in Python; for what a long recording costs in speech-to-text, see 36 lectures, three chunks each.
Sources
Related posts
More in Pricing
- Ideogram 4.5 blended cost per image: a 70/20/10 quality mix on Sume
A 70/20/10 low, medium, high mix of Ideogram 4.5 averages $0.0688 an image on Sume, $0.0725 with the fee. Three mixes and a 1,000-image bill.
- Ideogram 4.5: low drafts then one high final, or high only? Break-even
At $0.0396 a low draft and $0.2901 a high final all-in, six drafts plus one final cost $0.53, 1.8 high renders. When drafting pays, with the math.
- Ideogram 4.5 with n=4 on Sume: one write, four times the price
A POST /v1/images with n=4 spends one request from the write budget but bills cost_usd x n. Worst-case table for low, medium and high quality.
- Is a 540x960 preview render cheaper? Not on Sume Timeline
Sume prices Timeline 1.0 per whole output minute with no resolution term in the docs. Use the free plan to check a Short instead of paying for a small preview.
Written by Sume