Max text length for Gemini and OpenAI TTS: the docs name none

The OpenAI TTS guide and the Gemini speech page I read state no input limit. Sume publishes 20,000 characters and 1,200 s; here is a splitter.

5 min readSume
All posts

If you search for the maximum text length of OpenAI or Gemini text-to-speech, the pages I read give you no number. The OpenAI guide describes models, voices, formats and the instructions parameter but states no input limit, and the Gemini speech-generation page does not specify token limits. Sume's docs do publish one: 20,000 characters per transcript, with audio over 1,200 seconds failing as tts_duration_exceeded. A publishing gap is a planning risk, because you only learn the limit from a failed request.

Why an unpublished limit hurts

A limit you cannot read has to be found by trial. That means a production script discovers it at the worst time, usually on your longest chapter. Splitting defensively at a small size works, but you lose pacing continuity at every seam. A published number lets you split once, at the largest safe size, along sentence ends.

Input limits on the pages read, 2026-10-05
ServiceInput limit statedPage
OpenAI TTS guideNot stated on the pagedevelopers.openai.com text-to-speech guide
Gemini speech generationToken limits not specifiedai.google.dev speech generation guide
Sume TTS 1.0 and router1-20,000 characters; audio over 1,200 s failsSume repository docs and price catalog
Sume cost at the cap$0.95 for 20,000 characters20,000 x 0.00475 = 95 cents

A splitter that respects the Sume cap

Split on sentence ends, fill each chunk up to 20,000 characters, and estimate cost with the same ceil-to-cents rule the reserve uses. Spaces and punctuation count towards the limit and the price. The function below does that in plain Python.

import re

def split(text, cap=20_000):
    sentences = re.split(r"(?<=[.!?])\s+", text.strip())
    chunks, cur = [], ""
    for s in sentences:
        if len(s) > cap:
            raise ValueError("one sentence exceeds the cap")
        if len(cur) + len(s) + 1 > cap and cur:
            chunks.append(cur)
            cur = s
        else:
            cur = f"{cur} {s}".strip()
    return chunks + ([cur] if cur else [])

def cents(n):  # $0.0475 per 1,000 characters, ceil, minimum 1 cent
    return max(1, -(-n * 475 // 100_000))

for c in split(open("script.txt").read()):
    print(len(c), cents(len(c)))

Checking the audio side

The character cap is not the only limit. A slow generation_config.speed of 0.6 stretches the same text, and a chunk that runs past 1,200 seconds fails without being captured. For a very long chunk at low speed, split earlier than the cap. A failed job is not captured, but it still costs you a retry, so a safer chunk of 12,000 characters is a fair default for slow reads.

Use one Idempotency-Key per chunk so a retry returns the existing job instead of billing twice.

Fill in the table with a probe, not a guess

If you must run on a service with no published limit, probe it once in a test project: send 1,000, 4,000 and 10,000 characters of dull text, record what fails and how, and put the result in your own runbook with the date. Re-run it when the vendor ships a new model, because a limit is a property of a model version. Write the result next to the vendor's page URL so a teammate can see what was documented and what you measured.

For Sume you do not need a probe: the cap is in the docs, the catalog row reports capabilities.max_characters, and a request over the limit is rejected before it is queued or charged.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume