Gemini 3.5 Transcribe API limits: 1 hour vs Sume STT's 10 minutes
Gemini 3.5 Transcribe takes up to 1 hour per request, 30 minutes with diarization or word timestamps. Sume STT takes 600 seconds; split longer audio.

Gemini 3.5 Transcribe accepts up to 1 hour of audio per request, dropping to 30 minutes when speaker diarization or word-level timestamps are enabled. Sume STT (sume/stt-1.0) takes one audio file per job with a documented maximum of 10 minutes, so anything longer has to be split before you submit it.
Gemini limits are from its model page and the changelog entry of August 26, 2026, which lists general availability; Sume's limit is from the duration_seconds field in the OpenAPI document. All read 2026-09-30.
What are the per-request limits side by side?
| Item | Gemini 3.5 Transcribe | Sume STT 1.0 |
|---|---|---|
| Plain transcription | Up to 1 hour | Maximum 10 minutes |
| With word timestamps | Up to 30 minutes | Always returned, no flag |
| With diarization | Up to 30 minutes | Not exposed |
| Input | Audio file | Public HTTPS audio_url |
| Languages | 85+ auto-detected | Auto-detect, or a language_code hint |
What does duration_seconds do on Sume?
duration_seconds is an integer from 1 to 600 that the docs describe as the audio duration used for usage reservation. Omit it and Sume reserves for 1 minute. Set it to the real length of each chunk so the reservation matches the audio you send.
How do I handle audio longer than 10 minutes?
Cut the recording into pieces of 600 seconds or less, submit each as its own job, and add each piece's offset to the word start and end values when you stitch the transcripts back together. Cut at silence rather than mid-sentence. Timeline audio can slice Sume-hosted audio into ranges (operation: split); see split an audio file into parts, and the whole flow is in transcribe long audio files.
Which limit should decide my choice?
If your files are one-hour recordings and you want them in one call, the Gemini limit fits and the Sume limit does not without splitting. If your audio is already in clips under 10 minutes, the cap does not matter. Word timings come back on every Sume job, so the 30-minute reduction that Gemini applies for timestamps has no Sume equivalent.
Sources
Related posts
More in Models
- Gemini API speaker diarization: 8 speakers; Sume STT has no labels
Gemini 3.5 Transcribe diarizes up to 8 speakers, 3+ experimental, within 30 minutes per request. Sume STT returns word timings but no speaker labels.
- Gemini Omni video editing: the EEA, UK and under-10-second rules
Google says editing or extending uploaded videos with Omni is unavailable in the EEA, Switzerland and the UK, and uploads must be under 10 seconds.
- Gemini Omni extend video: Google's 40 s cap vs Sume's modes
Google's Gemini Omni can extend a video by 3-10 s up to 40 s. Sume's Omni row has no extend mode; here is what it offers instead.
- Gemini TTS API: which engine Sume's TTS routes to
Gemini 3.8 TTS is not in Sume's TTS Router, which lists Cartesia Sonic ids only. What Sume's TTS takes for voice, engine and length, with a checklist.
Written by Sume