Podcast sponsor ad reads: 36 spots of 450 characters, $1.08 on Sume

36 podcast ad reads of 450 characters cost $1.08 on Sume TTS at 3 cents each. MAI-Voice-2.1 lists $0.36 and Flash $0.24 for the same characters.

5 min readSume
All posts

Thirty-six sponsor reads of 450 characters each cost $1.08 on Sume TTS 1.0, because each read is its own job and 450 characters is 2.14 cents, which bills as 3 cents. The same 16,200 characters at Microsoft's listed prices are $0.36 for MAI-Voice-2.1 ($22 per 1M characters) and $0.24 for MAI-Voice-2.1-Flash ($15 per 1M), read 2026-10-08.

MAI-Voice is cheaper per character. Sume's price is a published per-job rounding on top of $0.0475 per 1,000 characters, and the cents add up when you render many short reads.

How a 450-character read is billed

Sume TTS 1.0 charges $0.0475 per 1,000 characters, with a one-cent minimum per job. 450 characters is 450 x 0.00475 = 2.1375 cents, and the billed amount rounds up to 3 cents. Spaces and punctuation count. Assuming about 150 spoken words a minute and six characters per word with the space, 450 characters is roughly 75 words, or 30 seconds, a normal host-read spot length. Those speaking-rate figures are assumptions, not vendor numbers.

The 36 reads here are 12 sponsors with three variants each. Using one job per read keeps retakes cheap: redoing a single read is 3 cents.

36 ad reads of 450 characters, 16,200 characters total (read 2026-10-08)
OptionRateTotal
Sume TTS 1.0, 36 jobs$0.0475 per 1,000 characters, rounded up per job$1.08
MAI-Voice-2.1$22 per 1M characters$0.36
MAI-Voice-2.1-Flash$15 per 1M characters$0.24

Does a read fit?

Microsoft's launch post gives Flash a 45-second audio limit; 450 characters at the speaking rate above is about 30 seconds, so it fits with margin. Sume allows up to 20,000 characters and 1,200 seconds of audio per request, so length is not the constraint here.

Sume's output is MP3 at 44.1 kHz and 128 kbps by default; WAV is available through output_format. Tone is set with generation_config (speed 0.6 to 1.5, volume 0.5 to 2.0, emotion as free text).

Render one read

Each sponsor read is an asynchronous job. Use a distinct Idempotency-Key per sponsor and variant so a retry never double-bills. The voice id comes from your workspace; YOUR_VOICE_ID below is a placeholder.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sponsor-07-variant-b" \
  -d '{"transcript": "This episode is brought to you by Acme. Use code POD for 20 percent off.",
       "voice": {"mode": "id", "id": "YOUR_VOICE_ID"},
       "generation_config": {"speed": 0.95},
       "mode": "async"}'

Disclosure and review

Ad reads in an AI voice should be disclosed in the way your platform and the sponsor require, and a sponsor will usually want to approve the audio, not just the text. Keep every take with its job id and script so that you can show what was approved. Because each read is a separate cheap job, regenerate a line when the sponsor changes a word rather than splicing audio by hand.

Check numbers, promo codes and URLs by ear or with a transcript check. A read that says the wrong code costs the sponsor more than the whole batch cost to generate.

Sources

Related posts

More in Pricing

All Pricing posts

Written by Sume