20-minute Reel narration vs Sume's 20,000-character TTS limit

Instagram allows Reels up to 20 minutes. Sume TTS accepts 1 to 20,000 characters per request. A rough estimate of how much narration one request covers.

4 min readSume
All posts

A single Sume text-to-speech request takes 1 to 20,000 characters of transcript, and at a typical narration pace that is somewhere around 20 minutes of speech, so one request is close to the whole of a maximum-length Reel. That second number is my estimate, not a Sume figure: it assumes about six characters per word including the space and roughly 150 spoken words per minute, both of which vary by language and voice. Instagram's Reels page says Reels can be up to 20 minutes (read 2026-10-03). The character limit comes from the TTS Router and TTS 1.0 API notes in the Sume repository, which say the router follows the same 1-20000 character rule.

The arithmetic

Treat these as planning numbers. Voices, punctuation and the language change the real duration, so measure a sample and use the actual result.

Estimated narration length for 20,000 characters, read 2026-10-03
Assumed paceWords (6 chars each)Minutes
120 wpm3,33327.8
150 wpm3,33322.2
180 wpm3,33318.5

Measure instead of guess

The most reliable approach is to count your script's characters and split it before it gets near the limit. The function below splits on sentence ends so that no chunk exceeds a budget, and prints the chunk count and the rough minutes using the same assumptions as the table. Keep a margin under 20,000 so a revision does not push a chunk over.

import re

LIMIT = 20000
CHARS_PER_WORD = 6
WPM = 150

def chunks(script, budget=18000):
    sentences = re.split(r"(?<=[.!?])\s+", script.strip())
    out, current = [], ""
    for sentence in sentences:
        if current and len(current) + len(sentence) + 1 > budget:
            out.append(current)
            current = sentence
        else:
            current = f"{current} {sentence}".strip()
    if current:
        out.append(current)
    return out

script = "This is a sentence about the product. " * 600
parts = chunks(script)
for number, part in enumerate(parts, start=1):
    minutes = len(part) / CHARS_PER_WORD / WPM
    print(number, len(part), "chars", round(minutes, 1), "min (estimate)")
assert all(1 <= len(part) <= LIMIT for part in parts)

If you need more than one request

Splitting into several requests gives you several audio files, and you then need to join them. Sume documents a timeline audio job that concatenates up to 20 Sume-hosted parts into one gapless file for $0.01, and parts must share one channel layout. So narration beyond one request is a joining problem rather than a wall.

Whether a 20-minute narrated Reel is a good idea is a separate question. Instagram's page also says videos longer than 3 minutes may not be recommended to new audiences, so a long narrated Reel suits an existing audience more than discovery. Check the TTS Router catalog for the model and the price per character, because the price depends on the model and includes Sume's margin.

Planning a long narrated Reel

Also decide where the pauses go. Sentence boundaries in the script are natural places for a breath, and the synthesized audio follows your punctuation, so a script with short sentences produces a different total length than the same words in long ones. Keep a note of the measured seconds per thousand characters for each voice you use, and the next plan becomes a multiplication.

Write the script in sections that map to scenes, and keep each section well under the character limit. That way each request produces one scene's audio, you can regenerate a single section without touching the others, and the joined file can be re-based against your video slots using the offsets the concat job returns.

Sume's timeline audio concat returns the part offsets so you can align Timeline video starts to them, which is the piece that keeps long narration and picture in step. Without it, you would be estimating where each section begins.

Check the voice and language before you write the full script. Some languages need more characters per spoken second than English, which pulls the real duration away from the table above. Generate a sample of one section, measure its length with a probe, and recompute the plan from the measured rate instead of the assumed 150 words per minute.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume