Gemini TTS limit: 16,384 output tokens at 25 per second is 10:55

Gemini 3.8 TTS is reported to cap output at 16,384 tokens with 25 audio tokens per second, about 10 minutes 55 seconds. Sume's TTS cap is 1,200 seconds.

4 min readSume
All posts

Digital Applied's September 2026 tracker reports a 16,384-token output limit for Gemini 3.8 TTS and 25 audio tokens per second. Divide the two: 16,384 / 25 is 655.36 seconds, which is 10 minutes 55 seconds of speech in one response. That is arithmetic on a third-party report, not a figure from Google's page, so confirm the limit in your own project before you plan a long read around it.

The Gemini API changelog (read 2026-10-02) confirms the model reached general availability on 2026-09-22. For Sume, the number to compare is its own cap: synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured.

How do the two limits compare?

Both are per request, and both are solved the same way: split the script. The table puts the figures side by side; the Gemini row is reported by Digital Applied.

Per-request audio limits, read 2026-10-02.
SurfaceLimitSource
Gemini 3.8 TTS output16,384 tokens, 25 tokens per second, about 655 sReported by Digital Applied
Sume TTS 1.0 synthesized audio1,200 s; longer fails with tts_duration_exceededSume OpenAPI
Sume TTS transcript20,000 characters per requestSume OpenAPI

How do I split a long script?

Cut at sentence or paragraph boundaries, submit each piece as its own job, and keep the job ids in order. Sume's TTS can return word timings and sentence segments, so you can also check where each piece ends. Use a stable voice for every piece so the seams do not change timbre.

How do I join the pieces?

Sume's timeline audio concat takes 1 to 20 ordered parts from your workspace's media.sume.com audio and joins them in the sample domain, with no re-synthesis and no silence at the seams. The result returns one audio_url, duration_seconds, and segments with start offsets. Twenty parts of up to ten minutes is well over three hours of narration.

Should I submit long reads as async jobs?

Yes. The sync wait is bounded at 30 seconds and bounds the HTTP wait, not the job. Submit with async, store the job id, and poll status_url, as the jobs page says; do not resubmit a paid request because a local wait timed out.

What does 655 seconds mean for a real script?

Speaking rate varies by voice and language, so translate seconds into words with your own sample rather than a rule of thumb. Generate one paragraph, measure the audio length, and divide the word count by the seconds. Multiply that rate by 655 for the largest script one Gemini response could hold, per the reported figures, and by 1,200 for the largest script one Sume job can hold.

Plan for margins. A script cut right at the limit is fragile, because a slower voice or a different language changes the length. Cut at about 80 percent of the limit and you will rarely meet the error.

Is the 655 second figure a promise?

No. It comes from two reported numbers on a third-party page, and Google's own page, read on 2026-10-02, says only that the model is generally available. Treat 655 seconds as an estimate and confirm the limit with a test request in your project before you build a pipeline around it.

Sources

Related posts

More in Models

All Models posts

Written by Sume