Gemini 3.8 Live audio: wrap 24 kHz PCM in WAV, resample to 16 kHz
Gemini 3.8 Live takes 16-bit 16 kHz PCM in and returns 24 kHz out. A Python WAV wrapper, an ffmpeg resample command, and the Sume detach settings that match.

The Gemini Live docs specify input as 16-bit PCM at 16 kHz and output at 24 kHz, over a WebSocket. So the audio you send must be 16 kHz mono PCM, and the raw audio you get back is 24 kHz PCM that needs a WAV header before most tools will play it.
Below are both conversions: a short Python function to wrap received PCM in a WAV file, and an ffmpeg command to produce the 16 kHz input.
The audio formats involved
Facts from the Live and speech generation pages, read 2026-10-03.
| Direction | Format | Source page |
|---|---|---|
| Live input | 16-bit PCM, 16 kHz | Live API |
| Live output | 24 kHz | Live API |
| Unary TTS output | WAV, 24 kHz, 16-bit PCM | Speech generation |
| Streaming TTS output | Raw L16, 24 kHz | Speech generation |
| Live transport | WebSocket (WSS) | Live API |
Wrap raw PCM in a WAV container
Raw L16 has no header, so a player cannot tell its rate. Python's standard wave module writes the header. The channel count is a parameter because the pages quoted here give rate and bit depth, not channels, so set it from what you receive.
import wave
def pcm_to_wav(pcm: bytes, path: str, rate: int = 24000, channels: int = 1):
with wave.open(path, "wb") as w:
w.setnchannels(channels)
w.setsampwidth(2) # 16-bit
w.setframerate(rate)
w.writeframes(pcm)
# usage: pcm_to_wav(b"".join(chunks), "reply.wav")Produce 16 kHz mono input
Run this once on any source file to get 16-bit PCM at 16 kHz, single channel, as a WAV.
ffmpeg -i input.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le voice-in.wavIf the source is a video on Sume
For audio that lives in a Sume-hosted video, skip the local step.
POST /v1/audio-detachwithsample_rate: 16000andchannels: "mono". The Sume docs call that combination the speech-to-text shape.- The default format is WAV (
pcm_s16le); the job is $0.01 and needs anIdempotency-Key. - Output over 900 seconds needs a
range, since output is capped at 900 s. - Poll
GET /v1/jobs/:id/status, then read the audio URL fromGET /v1/jobs/:id/result.
Where Sume stops
Live sessions are a WebSocket you hold open. The Sume Developer API has no WebSocket or SSE transport, so a voice agent that needs a generation job starts it by HTTP from your server and picks up the result by polling or webhook. See the browser voice app post for the key-handling side.
Sources
Related posts
More in Developers
- Gemini CLI v0.63 plan execution in CI: gate paid Sume calls first
Gemini CLI preview v0.63.0 adds autonomous plan execution in non-interactive mode. Before unattended runs, gate Sume paid tools with dry_run and max_spend_usd.
- Google's June 15 deprecation notice gave 15 and 63 days: run a drill
Google announced Veo and Imagen 4 deprecations on Jun 15, 2026 with shutdowns Jun 30 and Aug 17. Here is a five-step drill that fits inside the shorter window.
- Grok Imagine video API: 15 s, 5 references, request-ID polling
xAI's video guide for grok-imagine-video-1.5 lists up to 15 seconds, up to 5 reference images and async polling by request ID. The same loop on Sume jobs.
- How long AI video vendors keep your file: Veo 2 days, Higgsfield 7+
Veo keeps videos two days, Higgsfield files at least seven, Sora Batch outputs were kept 24 hours. Retention facts and a download-on-complete script.
Written by Sume