Docs page to explainer video: chunk the script under the TTS cap

Split a long document into voice-over chunks under Sume's 20,000-character TTS request limit and price them at $0.0475 per 1,000 characters. Runnable Python.

3 min readSume
All posts

For a doc-to-video pipeline, split the text on paragraph boundaries into chunks under 20,000 characters, send each to text to speech, and join the audio with Timeline audio concat. Narration costs $0.0475 per 1,000 characters on Sume, so a 12,000-character page is $0.57.

The limit

Sume's text to speech bills per transcript character, spaces and punctuation included, with a maximum of 20,000 characters per request. Chunk on paragraphs, not mid-sentence, so each take ends where a breath belongs.

A chunker that prices itself

No network is needed; this prints the chunk count and the voice cost for any text.

LIMIT = 20000
RATE = 0.0475 / 1000

def chunks(text, limit=LIMIT):
    out, cur = [], ""
    for para in text.split("\n\n"):
        if cur and len(cur) + len(para) + 2 > limit:
            out.append(cur)
            cur = ""
        cur += ("\n\n" if cur else "") + para
    if cur:
        out.append(cur)
    return out

doc = ("A paragraph about setup. " * 40 + "\n\n") * 40
parts = chunks(doc)
print(len(doc), "chars,", len(parts), "chunks, $%.4f" % (len(doc) * RATE))

Price by document length

Voice at $0.0475 per 1,000 characters, read 2026-10-08
CharactersTTS requestsVoice cost
2,0001$0.095
6,0001$0.285
12,0001$0.57
20,0001$0.95
40,0002$1.90

Joining the takes

POST /v1/timeline-1.0/audio with operation: "concat" joins up to 20 parts into one spine at $0.01 per job, and returns segment offsets you use to re-base your video slots. Then pass that file as audio.url on the render. A merged voice file over 1,800 seconds is not accepted by the render, so split very long docs into several videos. The render is billed at $0.10 per output minute (rounded up), so a 10-minute explainer adds $1.00 on top of the voice.

From text to scenes

Chunks for narration are not the same as scenes. After the voice is done, probe its length, divide it into scenes where the topic changes, and give each scene a still or a short clip. Screenshots of the real product are the best visuals, since they cannot misrepresent the interface. Import them as workspace files and use them as still slots. Check any code or command you show against the page text, because an image model is not a reliable renderer of exact strings.

Related posts

More in Developers

All Developers posts

Written by Sume