Gemini TTS streams raw PCM; Sume returns a finished wav or raw file
Gemini 3.8 Flash TTS streams headerless audio/l16 at 24 kHz. Sume returns a hosted file from an async job. Wrap the PCM in WAV and match the formats.

Gemini 3.8 Flash TTS has two output modes: a unary call that returns a WAV (audio/wav, 24 kHz, 16-bit mono), and a streaming call whose default is headerless raw PCM (audio/l16, 24 kHz, 16-bit mono) (read 2026-10-07). A raw stream will not play in most players until you add a header. Sume returns neither: you get a finished, Sume-hosted file from a job, and you choose mp3, wav or raw in output_format.
What a Gemini stream gives you
The Gemini docs also list mu-law (audio/mulaw) and A-law (audio/alaw) and configurable rates of 24000, 16000 and 8000 Hz. A raw L16 stream is just samples, with no sample rate or channel count in the bytes, so you carry those facts yourself.
| Mode | Default format | Header |
|---|---|---|
| Unary | audio/wav, 24 kHz, 16-bit mono | Yes |
| Streaming | audio/l16, 24 kHz, 16-bit mono | No (raw PCM) |
| Optional | audio/mulaw, audio/alaw | Rates 24000, 16000, 8000 Hz |
Wrap the stream in a WAV
If you collected the stream bytes, Python's standard library can add the header. This is standalone and needs no API key.
import wave
def pcm_to_wav(pcm: bytes, path: str, rate: int = 24000) -> None:
with wave.open(path, "wb") as w:
w.setnchannels(1)
w.setsampwidth(2) # 16-bit samples
w.setframerate(rate)
w.writeframes(pcm)
if __name__ == "__main__":
silence = b"\x00\x00" * 24000 # one second of silence
pcm_to_wav(silence, "out.wav")
print("wrote out.wav")What Sume returns
A Sume TTS 1.0 job is async. The result carries a mirrored audio artifact, so there is no stream to reassemble. output_format.container accepts mp3 (the default, 44,100 Hz at 128 kbps), wav and raw; sample_rate accepts 8000, 16000, 22050, 24000, 44100 and 48000; encoding accepts pcm_f32le, pcm_s16le, pcm_mulaw and pcm_alaw.
To match Gemini's unary shape, ask for wav, pcm_s16le at 24000. The mono layout is not a field in the Sume schema, so check the file you get back before you rely on a channel count.
Pick by what you do with the audio
Sume does not call Gemini TTS, so converting between the two is your pipeline's job, not a Sume option.
- Play it as it is generated: Gemini streaming, plus your own header and buffering.
- Hand a file to a video render, a CDN or an avatar clip: a Sume job, then the hosted URL.
- Need exact sentence cuts: Sume with
wavorrawandsegmentation.mode: sentence; mp3 returns timings only.
Sources
Related posts
More in Developers
- Omni Flash 4K, 10 seconds: a webhook-mode /v1/videos request body
One curl body for gemini-omni-flash-1.1 at 4K, 16:9, 10 seconds with audio and a callback_url, plus what the docs say the submit reserves and what it cannot.
- Gemini Omni Flash edit on Sume: a Python preflight for bad fields
A 29-line Python preflight for Sume video-router edit requests on gemini-omni-flash-1.1: catches the fields the docs say it rejects before you pay for a job.
- Two failure channels in the Sume SDK: submit error vs failed job
A generateVideoV1 error means no job exists; a failed job means one did and billing was settled. Handle both channels in TypeScript without double-submitting.
- generation_spend_cap_usd on a Format run: null is $500, 0 is a 400
On a Format run request, omit generation_spend_cap_usd for the Format cap, send a number up to 500, null for the $500 maximum. 0 or above 500 returns 400.
Written by Sume