Narrate a 2,000-word blog post with an AI voice: cost and steps

A 2,000-word post is roughly 12,000 characters, so one Sume TTS job at $0.0475 per 1,000 characters. The steps, settings and what to check before publishing.

5 min readSume
All posts

A 2,000-word blog post is about 12,000 characters, which fits in one Sume text-to-speech request (limit 20,000 characters), costs about 12 x $0.0475 = $0.57 at the catalog price, and comes back as an mp3 you host yourself. The characters-per-word figure is an estimate for English prose, so count your actual text before you budget.

Limits and prices come from the Sume API reference and the catalog. Check the live catalog for the current price before a large batch.

How do you count the characters?

Sume bills per 1,000 characters of the transcript, and spaces and punctuation count. English prose averages roughly 6 characters per word including the space, so 2,000 words lands near 12,000, but code blocks, link text and headings change that. Paste the exact text you will send into a counter. Anything over 20,000 characters has to be split into more than one request.

The request cap is 20,000 characters, so a 12,000 character post fits in one job.

How should you prepare the text?

A blog post written for the eye does not always work for the ear. Do these edits before you submit:

  • Replace code blocks, tables and raw URLs with a sentence that describes them.
  • Spell out symbols, abbreviations and numbers you want read a specific way.
  • Remove 'click here' and 'see below' phrases that make no sense in audio.
  • Add a short intro line with the title so listeners know what they are hearing.
  • Set language for non-English text; a voice-language mismatch is refused with 409 before any charge.

What does the request look like?

One request, one voice, one model. Use a wav container only if you will edit the audio again; the default mp3 at 44.1 kHz and 128 kbps is fine for publishing. If you want slower or faster delivery, generation_config.speed accepts 0.6 to 1.5.

curl -X POST https://api.sume.com/v1/tts-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: narration-001" \
  -d '{
    "model": "sonic-3.6",
    "transcript": "Welcome back. Today we compare two ways to get a voiceover.",
    "voice": { "id": "'"$VOICE_ID"'" },
    "language": "en"
  }'

Poll the job and download the artifact:

curl https://api.sume.com/v1/jobs/$JOB_ID/status \
  -H "Authorization: Bearer $SUME_API_KEY"

curl https://api.sume.com/v1/jobs/$JOB_ID/result \
  -H "Authorization: Bearer $SUME_API_KEY"

How do you add a sound bed or an intro?

Sume's TTS does not add music. If you want an intro sting or a quiet bed, generate it with the Music Router for $0.125, then join audio with timeline audio concat ($0.01 per job) or place the voice and the bed in a Timeline render. Both inputs must be Sume-hosted files.

If you want captions or a read-along highlight, request timestamps.words: true on the TTS job to get word timings in the result.

Cost of narrating a 2,000-word post, read 2026-10-02
ItemQuantityCatalog priceEstimate
SpeechAbout 12,000 characters$0.0475 per 1,000About $0.57
Intro bed (optional)1 music generation$0.125 each$0.125
Join (optional)1 concat job$0.01 flat$0.01
Retake of one sectionCharacters in that section$0.0475 per 1,000Pennies

What should you check before publishing?

Listen to the whole file once, with the text in front of you. Names, brand terms and numbers are where voices slip. If one paragraph is wrong, regenerate only that paragraph and join it back, rather than redoing 12,000 characters. Label the audio as AI-generated if the platform or your audience expects it, and keep the job id with the post so you can reproduce the take.

A worked example makes the budget concrete. Say the post is 11,800 characters after cleanup. One request covers it, so there is one job and one result. If you also want a 20-second intro bed and a join, add the music generation and one concat job. The total stays under a dollar, and most of the real cost is your time listening, not the API bill. If you publish weekly, keep the voice id, the model id and the language in one config file so every episode sounds the same.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume