Docs page to explainer video: chunk the script under the TTS cap
Split a long document into voice-over chunks under Sume's 20,000-character TTS request limit and price them at $0.0475 per 1,000 characters. Runnable Python.

For a doc-to-video pipeline, split the text on paragraph boundaries into chunks under 20,000 characters, send each to text to speech, and join the audio with Timeline audio concat. Narration costs $0.0475 per 1,000 characters on Sume, so a 12,000-character page is $0.57.
The limit
Sume's text to speech bills per transcript character, spaces and punctuation included, with a maximum of 20,000 characters per request. Chunk on paragraphs, not mid-sentence, so each take ends where a breath belongs.
A chunker that prices itself
No network is needed; this prints the chunk count and the voice cost for any text.
LIMIT = 20000
RATE = 0.0475 / 1000
def chunks(text, limit=LIMIT):
out, cur = [], ""
for para in text.split("\n\n"):
if cur and len(cur) + len(para) + 2 > limit:
out.append(cur)
cur = ""
cur += ("\n\n" if cur else "") + para
if cur:
out.append(cur)
return out
doc = ("A paragraph about setup. " * 40 + "\n\n") * 40
parts = chunks(doc)
print(len(doc), "chars,", len(parts), "chunks, $%.4f" % (len(doc) * RATE))Price by document length
| Characters | TTS requests | Voice cost |
|---|---|---|
| 2,000 | 1 | $0.095 |
| 6,000 | 1 | $0.285 |
| 12,000 | 1 | $0.57 |
| 20,000 | 1 | $0.95 |
| 40,000 | 2 | $1.90 |
Joining the takes
POST /v1/timeline-1.0/audio with operation: "concat" joins up to 20 parts into one spine at $0.01 per job, and returns segment offsets you use to re-base your video slots. Then pass that file as audio.url on the render. A merged voice file over 1,800 seconds is not accepted by the render, so split very long docs into several videos. The render is billed at $0.10 per output minute (rounded up), so a 10-minute explainer adds $1.00 on top of the voice.
From text to scenes
Chunks for narration are not the same as scenes. After the voice is done, probe its length, divide it into scenes where the topic changes, and give each scene a still or a short clip. Screenshots of the real product are the best visuals, since they cannot misrepresent the interface. Import them as workspace files and use them as still slots. Check any code or command you show against the page text, because an image model is not a reliable renderer of exact strings.
Related posts
More in Developers
- Does Seedance audio cost extra on Sume? generate_audio and the price
Seedance prices on Sume are per video token and carry no separate audio line. What generate_audio does, what it defaults to, and the 10 s price either way.
- Draft on Grok Imagine, finish on Seedance 2.5: two calls, one prompt
A Python flow that renders a prompt on grok-imagine-video-1.5 for $0.125, then re-sends the approved text to seedance-2.5 at 1080p for $14.22.
- Elixir Req: submit and poll a Sume video job after Sora
An Elixir script using Req: POST /v1/videos with an Idempotency-Key, poll every 30 seconds, stop on completed, failed or cancelled. Req never retries the POST.
- Env and secrets diff for removing Sora: SUME_API_KEY, one auth header
Which environment variables to delete, add and rotate when a service leaves the OpenAI Videos API for Sume, plus a startup check that fails on a missing secret.
Written by Sume