Gemini TTS limit: 16,384 output tokens at 25 per second is 10:55
Gemini 3.8 TTS is reported to cap output at 16,384 tokens with 25 audio tokens per second, about 10 minutes 55 seconds. Sume's TTS cap is 1,200 seconds.

Digital Applied's September 2026 tracker reports a 16,384-token output limit for Gemini 3.8 TTS and 25 audio tokens per second. Divide the two: 16,384 / 25 is 655.36 seconds, which is 10 minutes 55 seconds of speech in one response. That is arithmetic on a third-party report, not a figure from Google's page, so confirm the limit in your own project before you plan a long read around it.
The Gemini API changelog (read 2026-10-02) confirms the model reached general availability on 2026-09-22. For Sume, the number to compare is its own cap: synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and no credit is captured.
How do the two limits compare?
Both are per request, and both are solved the same way: split the script. The table puts the figures side by side; the Gemini row is reported by Digital Applied.
| Surface | Limit | Source |
|---|---|---|
| Gemini 3.8 TTS output | 16,384 tokens, 25 tokens per second, about 655 s | Reported by Digital Applied |
| Sume TTS 1.0 synthesized audio | 1,200 s; longer fails with tts_duration_exceeded | Sume OpenAPI |
| Sume TTS transcript | 20,000 characters per request | Sume OpenAPI |
How do I split a long script?
Cut at sentence or paragraph boundaries, submit each piece as its own job, and keep the job ids in order. Sume's TTS can return word timings and sentence segments, so you can also check where each piece ends. Use a stable voice for every piece so the seams do not change timbre.
How do I join the pieces?
Sume's timeline audio concat takes 1 to 20 ordered parts from your workspace's media.sume.com audio and joins them in the sample domain, with no re-synthesis and no silence at the seams. The result returns one audio_url, duration_seconds, and segments with start offsets. Twenty parts of up to ten minutes is well over three hours of narration.
Should I submit long reads as async jobs?
Yes. The sync wait is bounded at 30 seconds and bounds the HTTP wait, not the job. Submit with async, store the job id, and poll status_url, as the jobs page says; do not resubmit a paid request because a local wait timed out.
What does 655 seconds mean for a real script?
Speaking rate varies by voice and language, so translate seconds into words with your own sample rather than a rule of thumb. Generate one paragraph, measure the audio length, and divide the word count by the seconds. Multiply that rate by 655 for the largest script one Gemini response could hold, per the reported figures, and by 1,200 for the largest script one Sume job can hold.
Plan for margins. A script cut right at the limit is fragile, because a slower voice or a different language changes the length. Cut at about 80 percent of the limit and you will rarely meet the error.
Is the 655 second figure a promise?
No. It comes from two reported numbers on a third-party page, and Google's own page, read on 2026-10-02, says only that the model is generally available. Treat 655 seconds as an estimate and confirm the limit with a test request in your project before you build a pipeline around it.
Sources
Related posts
More in Models
- gpt-5.6-terra and gpt-5.6-luna on Sume: which still run, as what
Sume retired GPT 5.6 from every picker. Terra has no successor and runs as itself, Luna moves to GPT-6 Luna only where that is admitted, Sol moves to GPT-6 Sol.
- GPT-6.1 Sol effort on Sume: picker offers Low, Medium, High, plus Fast
OpenAI lists five effort values for GPT-6.1 Sol. Sume's picker exposes three, keeps effort off the model id, and prices Fast as an opt-in. What that means.
- GPT Image 2.5 edit drifts after a few turns: repeat the preserve list
OpenAI says to repeat the preserve list on each iteration to reduce drift. How to run a one-change-per-call edit chain on Sume, which keeps no chat memory.
- GPT Image 2.5 repeats the headline: say how many times it appears
Duplicate text in a GPT Image 2.5 result? OpenAI's guide says to quote the copy and state how often it appears. The prompt, plus n and billing on Sume.
Written by Sume