Narrate a 2,000-word blog post with an AI voice: cost and steps
A 2,000-word post is roughly 12,000 characters, so one Sume TTS job at $0.0475 per 1,000 characters. The steps, settings and what to check before publishing.

A 2,000-word blog post is about 12,000 characters, which fits in one Sume text-to-speech request (limit 20,000 characters), costs about 12 x $0.0475 = $0.57 at the catalog price, and comes back as an mp3 you host yourself. The characters-per-word figure is an estimate for English prose, so count your actual text before you budget.
Limits and prices come from the Sume API reference and the catalog. Check the live catalog for the current price before a large batch.
How do you count the characters?
Sume bills per 1,000 characters of the transcript, and spaces and punctuation count. English prose averages roughly 6 characters per word including the space, so 2,000 words lands near 12,000, but code blocks, link text and headings change that. Paste the exact text you will send into a counter. Anything over 20,000 characters has to be split into more than one request.
The request cap is 20,000 characters, so a 12,000 character post fits in one job.
How should you prepare the text?
A blog post written for the eye does not always work for the ear. Do these edits before you submit:
- Replace code blocks, tables and raw URLs with a sentence that describes them.
- Spell out symbols, abbreviations and numbers you want read a specific way.
- Remove 'click here' and 'see below' phrases that make no sense in audio.
- Add a short intro line with the title so listeners know what they are hearing.
- Set
languagefor non-English text; a voice-language mismatch is refused with 409 before any charge.
What does the request look like?
One request, one voice, one model. Use a wav container only if you will edit the audio again; the default mp3 at 44.1 kHz and 128 kbps is fine for publishing. If you want slower or faster delivery, generation_config.speed accepts 0.6 to 1.5.
curl -X POST https://api.sume.com/v1/tts-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: narration-001" \
-d '{
"model": "sonic-3.6",
"transcript": "Welcome back. Today we compare two ways to get a voiceover.",
"voice": { "id": "'"$VOICE_ID"'" },
"language": "en"
}'
Poll the job and download the artifact:
curl https://api.sume.com/v1/jobs/$JOB_ID/status \
-H "Authorization: Bearer $SUME_API_KEY"
curl https://api.sume.com/v1/jobs/$JOB_ID/result \
-H "Authorization: Bearer $SUME_API_KEY"
How do you add a sound bed or an intro?
Sume's TTS does not add music. If you want an intro sting or a quiet bed, generate it with the Music Router for $0.125, then join audio with timeline audio concat ($0.01 per job) or place the voice and the bed in a Timeline render. Both inputs must be Sume-hosted files.
If you want captions or a read-along highlight, request timestamps.words: true on the TTS job to get word timings in the result.
| Item | Quantity | Catalog price | Estimate |
|---|---|---|---|
| Speech | About 12,000 characters | $0.0475 per 1,000 | About $0.57 |
| Intro bed (optional) | 1 music generation | $0.125 each | $0.125 |
| Join (optional) | 1 concat job | $0.01 flat | $0.01 |
| Retake of one section | Characters in that section | $0.0475 per 1,000 | Pennies |
What should you check before publishing?
Listen to the whole file once, with the text in front of you. Names, brand terms and numbers are where voices slip. If one paragraph is wrong, regenerate only that paragraph and join it back, rather than redoing 12,000 characters. Label the audio as AI-generated if the platform or your audience expects it, and keep the job id with the post so you can reproduce the take.
A worked example makes the budget concrete. Say the post is 11,800 characters after cleanup. One request covers it, so there is one job and one result. If you also want a 20-second intro bed and a join, add the music generation and one concat job. The total stays under a dollar, and most of the real cost is your time listening, not the API bill. If you publish weekly, keep the voice id, the model id and the language in one config file so every episode sounds the same.
Sources
Related posts
More in Use cases
- New York FAIR News Act: AI video labels and what is pending
New York's FAIR News Act would require a label on news video substantially made by generative AI. It awaits the governor; here is what the sources say.
- New York synthetic performer ad law: dubbing and audio-only exceptions
New York S8420A exempts AI used only to translate a human performer's language, audio-only ads and expressive works. What that means for a dubbed ad.
- Burn a rolling headline sequence onto a muted news clip
Put timed headlines on a muted news clip with video-captions cues, a card colour and anchor_ratio. $0.20 per clip up to 60 seconds, no speech needed.
- News video reaches 77% weekly: cut platform-native versions
Reuters Institute's 2026 report says 77% watch online news video weekly. Cut one finished piece into vertical and landscape versions with Sume's video trim.
Written by Sume