Gemini TTS 24/16/8 kHz vs Sume TTS six sample rates up to 48 kHz

Gemini 3.8 Flash TTS offers 24000, 16000 and 8000 Hz. Sume TTS 1.0 accepts six rates from 8000 to 48000 Hz. Which rate suits video, phone and web.

4 min readSume
All posts

For video work, ask Sume for 48000 Hz or 44100 Hz. For phone playback, 8000 Hz. Gemini 3.8 Flash TTS lists configurable rates of 24000, 16000 and 8000 Hz (read 2026-10-07), so the top of its range sits below what most video editors expect. Sume's TTS 1.0 output_format.sample_rate accepts 8000, 16000, 22050, 24000, 44100 and 48000.

The two lists

Listing a rate does not mean the voice has content up there; a speech model that renders at 24 kHz and is resampled to 48 kHz fills the file size without adding detail. Check how your pipeline uses the file before you pay for the bigger one.

Sample rates available for TTS output - Gemini docs and Sume OpenAPI (read 2026-10-07)
RateGemini 3.8 Flash TTSSume TTS 1.0
8000 HzYesYes
16000 HzYesYes
22050 HzNot listedYes
24000 HzYes (default)Yes
44100 HzNot listedYes (default for mp3)
48000 HzNot listedYes

Matching the destination

  • Video editors and timelines: 44100 or 48000. Sume's own examples use wav, pcm_s16le at 44100 when feeding avatar mux.
  • Web playback: the mp3 default, 44100 Hz at 128 kbps, is a safe size.
  • Telephony: 8000 Hz with pcm_mulaw or pcm_alaw, both encodings in the schema.
  • Speech recognition pipelines: 16000 Hz is the usual input for transcribers.

Setting it on Sume

For mp3, the bit rate options are 32000, 64000, 96000, 128000 and 192000. For wav and raw, choose an encoding. Per-sentence slices exist only for wav and raw.

curl -X POST https://api.sume.com/v1/tts-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: rate-48k-001" \
  -d '{
    "transcript": "Ready for the editor.",
    "avatar_handle": "speaker",
    "language": "en",
    "output_format": {"container": "wav", "encoding": "pcm_s16le", "sample_rate": 48000}
  }'

A caution on file hand-offs

Sume Timeline audio takes only audio already on your workspace's media.sume.com, so generate the voice on Sume and pass the returned URL straight through. A file produced on another vendor must first be imported with POST /v1/media-imports. The Sume docs do not list Gemini TTS as an engine, so there is no format negotiation between the two inside Sume.

Sources

Related posts

More in Models

All Models posts

Written by Sume