Gemini TTS API polling: the Sume audio job pattern
Gemini 3.8 Flash TTS went GA on Sep 22. On Sume, speech is a job: submit async, store the job id, poll the status URL, and never resubmit while it runs.

The Gemini changelog for September 22, 2026 announces gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts as generally available, but the entry does not describe polling. For polling you need a job API. On Sume, text-to-speech is a job: submit with mode: "async", keep the job id, poll GET /v1/jobs/:id/status until terminal is true, then read the result. If a wait or your process ends early, poll again; do not resubmit.
Gemini facts are from its changelog; the job pattern is from Sume's Jobs and results, read 2026-10-01.
What did Gemini announce for TTS?
Two GA models: gemini-3.8-flash-tts, described as the flagship creative TTS model, and gemini-3.8-flash-lite-tts, described as cost-efficient and meant to replace gemini-3.1-flash-tts-preview. The same entry adds a Voices endpoint at /v1beta/voices. The entry is a release note, so how a request completes is not covered there.
How does polling work on Sume?
A submit is accepted once Sume has a durable job id, and every mode returns that id in the first response. A 2xx means the job exists and paid work is in flight, not that it finished. The docs tell you to store the id "so your integration can recover work after process restarts", and to use exponential backoff, stopping on completed, failed or canceled.
| Status | Terminal | Your next step |
|---|---|---|
queued | No | Keep polling |
processing | No | Keep polling |
completed | Yes | Read GET /v1/jobs/:id/result |
failed | Yes | Read the job record for the error |
canceled | Yes | Nothing to read |
const headers = { Authorization: "Bearer " + process.env.SUME_API_KEY };
async function waitForJob(jobId) {
let delayMs = 2000;
for (;;) {
const res = await fetch("https://api.sume.com/v1/jobs/" + jobId + "/status", { headers });
const s = await res.json();
if (s.terminal) return s;
delayMs = Math.min(delayMs * 2, 15000);
await new Promise((r) => setTimeout(r, (s.next_poll_after_seconds ?? delayMs / 1000) * 1000));
}
}Why not resubmit when a call times out?
Because the first submit already created a paid job. The docs say not to resubmit the original paid request just because a local process timed out. If the submit call itself failed and you do not know whether it landed, retry with the same Idempotency-Key so you get the original job back rather than a second one.
Where do voices and engines come in?
Voice and engine choice happen at submit time and do not change the polling loop. See finding a voice id on Sume for voices and which engine Sume routes for Gemini TTS for engine selection.
Sources
Related posts
More in Developers
- Gemini 3.8 TTS returns WAV by default; Sume TTS returns MP3
Gemini 3.8 TTS unary requests return WAV with a RIFF header. Sume TTS 1.0 defaults to mp3 at 44100 Hz, so ask for wav when your pipeline needs it.
- Gemini TTS reads stage directions aloud: Sume emotion field
Gemini 3.8 TTS speaks its text field verbatim, so style goes in speech_metadata. On Sume, keep the transcript clean and put delivery in generation_config.
- Gemini API file size limit: 2GB per file, 20GB storage
Gemini's rate-limits page lists a 2GB input file limit and 20GB file storage under Batch API limits. Sume takes public HTTPS URLs instead of uploaded files.
- Gemini Batch API rate limits vs Sume's 100-run bulk queue
Gemini Batch API allows 100 concurrent batch requests and a 2GB input file. Sume bulk runs queue up to 100 Format runs with concurrency 1 to 16.
Written by Sume