ElevenLabs v4 10,000-character limit vs Sume TTS 20,000 transcript

Higgsfield's changelog lists 10,000 characters per ElevenLabs v4 generation; Sume TTS 1.0 takes 20,000. How to split a long script.

4 min readSume
All posts

Higgsfield's changelog lists up to 10,000 characters per generation for ElevenLabs v4. Sume TTS 1.0's transcript field accepts up to 20,000 characters, but audio longer than 1,200 seconds fails with tts_duration_exceeded, so the shorter of the two limits wins in practice. At a typical pace of about 900 characters a minute (our assumption), 20,000 characters is around 22 minutes, which is already over the 1,200-second cap.

So split long scripts well before either limit.

Where do the limits sit?

The Higgsfield figure is a changelog entry for its integration of v4, not a vendor spec sheet; confirm it for the API you use.

Limits that decide a split. Higgsfield changelog and Sume OpenAPI, read 2026-10-02.
LimitValueSource
Characters per v4 generation10,000Higgsfield changelog
Sume TTS transcript20,000 charactersSume OpenAPI
Sume TTS audio length1,200 s, else tts_duration_exceededSume docs

How should I split a long script?

Split at paragraph breaks, not mid-sentence, so each part starts and ends on a natural pause. Keep each part under about 15 minutes of speech for Sume, and use the same voice id for every part. Generate parts as separate jobs and keep the job ids in order.

How do I join the parts?

Sume's timeline audio can concatenate audio into a reusable file, per the Timeline docs. Listen to every join for a change in pace or breath. If a join sounds wrong, regenerate only that part.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume