Gemini TTS 24/16/8 kHz vs Sume TTS six sample rates up to 48 kHz
Gemini 3.8 Flash TTS offers 24000, 16000 and 8000 Hz. Sume TTS 1.0 accepts six rates from 8000 to 48000 Hz. Which rate suits video, phone and web.

For video work, ask Sume for 48000 Hz or 44100 Hz. For phone playback, 8000 Hz. Gemini 3.8 Flash TTS lists configurable rates of 24000, 16000 and 8000 Hz (read 2026-10-07), so the top of its range sits below what most video editors expect. Sume's TTS 1.0 output_format.sample_rate accepts 8000, 16000, 22050, 24000, 44100 and 48000.
The two lists
Listing a rate does not mean the voice has content up there; a speech model that renders at 24 kHz and is resampled to 48 kHz fills the file size without adding detail. Check how your pipeline uses the file before you pay for the bigger one.
| Rate | Gemini 3.8 Flash TTS | Sume TTS 1.0 |
|---|---|---|
| 8000 Hz | Yes | Yes |
| 16000 Hz | Yes | Yes |
| 22050 Hz | Not listed | Yes |
| 24000 Hz | Yes (default) | Yes |
| 44100 Hz | Not listed | Yes (default for mp3) |
| 48000 Hz | Not listed | Yes |
Matching the destination
- Video editors and timelines: 44100 or 48000. Sume's own examples use wav,
pcm_s16leat 44100 when feeding avatar mux. - Web playback: the mp3 default, 44100 Hz at 128 kbps, is a safe size.
- Telephony: 8000 Hz with
pcm_mulaworpcm_alaw, both encodings in the schema. - Speech recognition pipelines: 16000 Hz is the usual input for transcribers.
Setting it on Sume
For mp3, the bit rate options are 32000, 64000, 96000, 128000 and 192000. For wav and raw, choose an encoding. Per-sentence slices exist only for wav and raw.
curl -X POST https://api.sume.com/v1/tts-1.0/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: rate-48k-001" \
-d '{
"transcript": "Ready for the editor.",
"avatar_handle": "speaker",
"language": "en",
"output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 48000}
}'A caution on file hand-offs
Sume Timeline audio takes only audio already on your workspace's media.sume.com, so generate the voice on Sume and pass the returned URL straight through. A file produced on another vendor must first be imported with POST /v1/media-imports. The Sume docs do not list Gemini TTS as an engine, so there is no format negotiation between the two inside Sume.
Sources
Related posts
More in Models
- Is Gemini 3.8 Flash TTS in the Live API? No, per Google's model page
Google's model page marks the Live API unsupported for Gemini 3.8 Flash TTS. For live voice use Gemini 3.8 Live; Sume TTS is for async audio jobs.
- Gemini TTS short pause and breath tags vs Sume sentence segments
Gemini 3.8 Flash TTS takes inline tags like short pause and breath. Sume lists no inline tags; here is how to control pacing with segments and speed.
- Gemini Omni Flash 1.1 API cost per clip, 360p to 4K
Gemini Omni Flash 1.1 on Sume is $0.0375 to $0.375 per second. An 8-second 720p clip is $1.00; 100 clips are $100. Full table by resolution.
- gemini-omni-flash-preview: Oct 22 shutdown on the table, not Sep 30
Google's release note says gemini-omni-flash-preview is deprecated on Sept 30, 2026, but its deprecations table now lists Oct 22. What to pin and what to test.
Written by Sume