Gemini 3.8 Live audio: wrap 24 kHz PCM in WAV, resample to 16 kHz

Gemini 3.8 Live takes 16-bit 16 kHz PCM in and returns 24 kHz out. A Python WAV wrapper, an ffmpeg resample command, and the Sume detach settings that match.

4 min readSume
All posts

The Gemini Live docs specify input as 16-bit PCM at 16 kHz and output at 24 kHz, over a WebSocket. So the audio you send must be 16 kHz mono PCM, and the raw audio you get back is 24 kHz PCM that needs a WAV header before most tools will play it.

Below are both conversions: a short Python function to wrap received PCM in a WAV file, and an ffmpeg command to produce the 16 kHz input.

The audio formats involved

Facts from the Live and speech generation pages, read 2026-10-03.

Gemini audio formats (read 2026-10-03)
DirectionFormatSource page
Live input16-bit PCM, 16 kHzLive API
Live output24 kHzLive API
Unary TTS outputWAV, 24 kHz, 16-bit PCMSpeech generation
Streaming TTS outputRaw L16, 24 kHzSpeech generation
Live transportWebSocket (WSS)Live API

Wrap raw PCM in a WAV container

Raw L16 has no header, so a player cannot tell its rate. Python's standard wave module writes the header. The channel count is a parameter because the pages quoted here give rate and bit depth, not channels, so set it from what you receive.

import wave

def pcm_to_wav(pcm: bytes, path: str, rate: int = 24000, channels: int = 1):
    with wave.open(path, "wb") as w:
        w.setnchannels(channels)
        w.setsampwidth(2)        # 16-bit
        w.setframerate(rate)
        w.writeframes(pcm)

# usage: pcm_to_wav(b"".join(chunks), "reply.wav")

Produce 16 kHz mono input

Run this once on any source file to get 16-bit PCM at 16 kHz, single channel, as a WAV.

ffmpeg -i input.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le voice-in.wav

If the source is a video on Sume

For audio that lives in a Sume-hosted video, skip the local step.

  • POST /v1/audio-detach with sample_rate: 16000 and channels: "mono". The Sume docs call that combination the speech-to-text shape.
  • The default format is WAV (pcm_s16le); the job is $0.01 and needs an Idempotency-Key.
  • Output over 900 seconds needs a range, since output is capped at 900 s.
  • Poll GET /v1/jobs/:id/status, then read the audio URL from GET /v1/jobs/:id/result.

Where Sume stops

Live sessions are a WebSocket you hold open. The Sume Developer API has no WebSocket or SSE transport, so a voice agent that needs a generation job starts it by HTTP from your server and picks up the result by polling or webhook. See the browser voice app post for the key-handling side.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume