3-minute Reel script with AI voice: characters, cost and the cap

A 180-second Reel narration is a character budget, not a word count. Sume TTS is $0.0475 per 1,000 characters with a 20,000-character cap; measure, then scale.

5 min readSume
All posts

A 3-minute Reel script for an AI voice is a character budget, and you find it by measuring one paragraph rather than by guessing a speaking rate. Sume TTS 1.0 is priced at $0.0475 per 1,000 transcript characters, spaces and punctuation included, with a maximum of 20,000 characters per request. Even a script that fills the whole cap costs $0.95, so the cost is not the constraint; the fit to 180 seconds is.

Instagram's Reels guide (read 2026-10-03) lists 60 seconds to 3 minutes as the range for longer storytelling, walkthroughs and detailed how-tos, and Metricool (read 2026-10-03) reports the 3-minute ceiling in its Instagram news page.

Measure before you write 180 seconds

Voices, languages and the speed setting all change how long a character count lasts, and Sume's docs do not publish a characters-per-second figure, so I will not either. Do this instead: render a 1,000-character sample in the voice you will use, read its duration_seconds, and divide. If the sample runs 62 seconds, the 180-second budget is about 2,900 characters at that voice and pace. generation_config.speed accepts 0.6 to 1.5 if you need to nudge the fit, but a rewrite usually sounds better than a fast read.

Character budget and cost

The table uses the catalog rate of $0.0475 per 1,000 characters. Cost is linear, so any other length is a multiplication.

TTS 1.0 cost by script length at $0.0475 per 1,000 characters (read 2026-10-03)
Script length (characters)CostNote
1,000$0.0475Sample to measure the voice
2,500$0.11875Typical short walkthrough
5,000$0.2375Fits one request
10,000$0.475Fits one request
20,000$0.95Hard cap for one request

Counting and splitting in code

This pure-Python helper counts characters the way the price does (spaces and punctuation included), prices the script, and splits it into chunks under 20,000 characters on sentence ends, so you never hit the per-request cap.

import re

RATE_PER_1000 = 0.0475
CAP = 20000

def chunks(script, cap=CAP):
    out, cur = [], ""
    for s in re.split(r"(?<=[.!?])\s+", script.strip()):
        if cur and len(cur) + 1 + len(s) > cap:
            out.append(cur)
            cur = s
        else:
            cur = (cur + " " + s).strip()
    if cur:
        out.append(cur)
    return out

def price(script):
    return round(len(script) / 1000 * RATE_PER_1000, 4)

if __name__ == "__main__":
    text = "Open on the problem. Show the fix. End on one clear step. " * 40
    print(len(text), price(text), len(chunks(text)))

Join the pieces and build the Reel

If you split the script, join the voice files with timeline audio: operation: "concat" takes 1 to 20 ordered parts, joins in the sample domain with no re-synthesis, costs $0.01 flat, and returns segments[] with each part's start offset, which you use as the start of the matching picture slots in Timeline 1.0. Timeline accepts an audio spine up to 1800 seconds, so a 180-second Reel is far inside it, and bills $0.10 per output minute rounded up: $0.30 for 180 seconds.

What Sume does not do

Sume does not write the script, count syllables or promise that a script will land at exactly 3:00. If the voice runs long, trim words or lift speed slightly and re-render, since a TTS job is cheap next to a video job. If the finished render is longer than 180 seconds, cut it before you post; a 181-second render bills as four Timeline minutes.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume