What does one hour of AI narration cost on Sume, MAI and ElevenLabs?

One hour of finished narration is about 45,000 characters: $2.16 on Sume in three jobs, against $0.99 on MAI-Voice-2.1. Rates read 2026-10-07.

4 min readSume
All posts

One hour of finished AI narration is roughly 45,000 characters of script, and on Sume that costs $2.16 when you split it into three jobs of 15,000 characters ($0.72 each), or $2.17 if you also join the three files into one with a Timeline audio concat. The same hour costs less on MAI-Voice-2.1 ($0.99 at the $22 per million characters rate) and more on ElevenLabs v3 ($3.60). Sume is not the cheapest line in the table below; it is the one where the arithmetic is flat and every number is published.

This post uses one assumption for every vendor so the comparison is fair: 750 characters of script per minute of speech. That figure comes from Cartesia's pricing page, which says one minute of Sonic audio takes 750 to 800 credits at one credit per character. Sume TTS 1.0 runs Sonic and meters one credit per character, so the figure carries over to Sume. Other vendors' voices may speak faster or slower, and a different pace moves every row by the same percentage.

The hour, priced per vendor

Sixty minutes at 750 characters a minute is 45,000 characters. Multiply by each vendor's published per-million-character rate:

Price of 45,000 characters (one hour at 750 characters a minute). Vendor rates read 2026-10-07; Sume rate from the TTS 1.0 catalog entry.
ServiceRate per 1M charactersOne hour
Sume TTS 1.0 (one rate, before per-job rounding)$47.50$2.1375
MAI-Voice-2.1$22$0.990
MAI-Voice-2.1-Flash$15$0.675
ElevenLabs Flash / Turbo$40$1.80
ElevenLabs v2 Multilingual and v3$80$3.60
OpenAI tts-1$15$0.675
OpenAI tts-1-hd$30$1.35

Why Sume bills three jobs, not one

A Sume TTS request accepts up to 20,000 characters, and a job fails with tts_duration_exceeded if the synthesized audio runs longer than 1,200 seconds. At 750 characters a minute, 1,200 seconds is 15,000 characters. So a 20,000-character request at a natural pace can overshoot the audio cap, and the safe chunk for an hour of narration is 15,000 characters or fewer. Three chunks of 15,000 characters cover the 45,000.

Each job is billed from its own character count: 15,000 characters at $0.0475 per 1,000 is $0.7125, which the catalog rounds up to $0.72. Three jobs come to $2.16. The single-rate figure in the table ($2.1375) is what you would pay with no rounding; the real bill sits a few cents above it.

Cut the script at paragraph breaks, not mid-sentence, so each chunk starts and ends cleanly. The three finished files can be joined with Timeline audio, which concatenates up to 20 parts at a flat $0.01 per job with no re-synthesis and no silence at the seams.

Submitting one chunk

Each chunk is one async job. Export VOICE_ID first, read the finished file from the job result, and keep a distinct idempotency key per chunk so a retry never bills twice.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hour-chunk-01" \
  -d @- <<JSON
{
  "transcript": "First 15,000 characters of the script go here.",
  "voice": { "id": "$VOICE_ID" },
  "language": "en",
  "output_format": { "container": "wav", "encoding": "pcm_s16le", "sample_rate": 44100 },
  "mode": "async"
}
JSON

Where the cheaper rows come with a catch

MAI-Voice-2.1 at $22 per million characters is less than half the Sume rate, and the Flash tier at $15 is lower still. Microsoft's own model page positions the standard model for audiobooks, content creation and voice-overs, and Flash for call centers, voice assistants and IVR. If you are already on Azure and the volume is large, the MAI line is a real saving.

OpenAI's gpt-4o-mini-tts is left out of the table because its audio output is billed in tokens ($12 per million audio output tokens, plus $0.60 per million input characters), and a token-to-minute conversion would be a guess. ElevenLabs lists v3 at 70+ languages and Flash at 32, which matters if your hour is not in English.

What Sume adds is the rest of the chain behind one key: the voice job, a music bed from Music Router at $0.125 per track, and an MP4 render at $0.10 per output minute. If your hour of narration needs only the audio file, compare the first column honestly. If it needs captions, a bed and a render, compare the whole bill.

Checks before you commit an hour

  • Run one 1,000-character sample and time it: if the audio is longer than 80 seconds, your pace is slower than 750 a minute and your hour needs more characters than 45,000.
  • Keep language set on every request for a non-English script; Sume's fallback inference only covers Hangul-only or kana-only text.
  • Store the job id for each chunk so a bad chunk can be redone alone without regenerating the rest.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume