Voicemaker 3,000-character cap vs a 20,000-character Sume TTS call

Voicemaker caps each conversion at 3,000 to 10,000 characters by plan; Sume TTS accepts up to 20,000 characters per request. What that means for long scripts.

5 min readSume
All posts

A 20,000-character script fits in one Sume TTS request, because the transcript field accepts 1 to 20,000 characters. On Voicemaker's pricing page (read 2026-10-10) a single conversion is capped at 3,000 characters on Starter, 5,000 on Creator and 10,000 on Pro, so the same script needs 7, 4 or 2 conversions.

That is the whole headline, and it is a workflow difference rather than a quality one. The rest of this post lays out the numbers from both sides so you can see where the split lands for your own script.

Voicemaker's limits and credits

Voicemaker sells monthly credit bundles with a per-conversion character ceiling. Its default voices cost 1 credit per character; its ProPlus expressive and high-resolution voices cost 4 credits per character. Credits also carry speech-to-text at 10 credits per second.

Voicemaker plan data from its pricing page (read 2026-10-10)
PlanPriceMonthly creditsPer-conversion limit
Free$0Limited250 characters
Starter$5200,000 (about 4 hours)3,000 characters
Creator$10500,000 (about 9 hours)5,000 characters
Pro$201,000,000 (about 18 hours)10,000 characters

The arithmetic for one long script

Take a 20,000-character narration. On default voices it uses 20,000 credits, which is 10% of Starter's 200,000 monthly credits. Starter's 3,000-character cap means 7 conversions (20,000 divided by 3,000 is 6.67, rounded up), and Pro's 10,000-character cap means 2.

On ProPlus voices at 4 credits per character the same script uses 80,000 credits, or 40% of Starter's allowance. The credit budget is fine; the per-conversion cap is the friction, because you must split the script and then join the audio yourself.

What a Sume TTS request carries

Sume TTS 1.0 takes a transcript of 1 to 20,000 characters and a voice selector. The selector is either an avatar (avatar_id or avatar_handle, found through GET /v1/avatar-1.0/avatars) or a voice.id that must be a voice UUID, not a name like alloy; the API rejects names so a bad id fails at request time.

Output defaults to mp3 at 44,100 Hz and 128 kbps, with wav and raw containers also available. generation_config accepts speed from 0.6 to 1.5 and volume from 0.5 to 2. Setting timestamps: { words: true } asks for word timings, and segmentation with mode sentence splits the result by sentence, with a boundary_lead_ms default of 70. Results are jobs, read through Jobs and results.

If a script is longer than 20,000 characters, split at a paragraph boundary and join the pieces with timeline audio, which concatenates Sume-hosted audio at the sample level with no re-synthesis.

Which limit hurts you

Both services have a ceiling; they just sit in different places.

  • Podcast chapter of 8,000 characters: one Pro conversion on Voicemaker, one request on Sume.
  • Audiobook chapter of 25,000 characters: three Voicemaker Pro conversions, or two Sume requests split near the middle.
  • Short ad of 600 characters: no limit matters on either side; compare voices instead.
  • Voice choice: Voicemaker's page lists voice cloning slots as a $2 per year add-on; Sume voices come from avatars or library ids.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume