Gemini 3.8 Flash TTS takes 8,192 input tokens: splitting a long script

Gemini 3.8 Flash TTS lists 8,192 input tokens and 16,384 output tokens. Sume TTS takes 20,000 characters per job. Here is how to split a long script.

4 min readSume
All posts

Gemini 3.8 Flash TTS lists an input limit of 8,192 tokens and an output limit of 16,384 tokens on its model page, so a long narration needs to be split into several requests. Sume TTS 1.0 accepts 1 to 20,000 characters per job and fails audio over 1,200 seconds.

The limits side by side

Older posts said Google named no text limit; the current model page now does.

TTS request limits, read 2026-10-07
ItemGemini 3.8 Flash TTSSume TTS 1.0
Input8,192 tokens1 to 20,000 characters
Output16,384 tokensAudio up to 1,200 s, else tts_duration_exceeded
UnitTokensCharacters

Split on sentence boundaries

Tokens are not characters, so do not guess a character cutoff for Gemini. Count tokens with Google's tooling and leave headroom. Split at sentence or paragraph ends so each clip starts and stops naturally, and keep the same voice and style text on every request.

Google treats the text as a verbatim transcript, per the speech generation docs, so keep style direction in the style field rather than the text, or it may be read aloud in every chunk.

On Sume

A 20,000-character ceiling covers most ad and explainer scripts in one job. Beyond that, split the script and join the clips with timeline audio, which concatenates 1 to 20 parts for a flat $0.01 per job (timeline audio docs). That is a published limit, not a promise about how seams sound: listen to every join.

Sume does not split text for you. You choose the break points.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume