Flow custom voices only work with ingredients: what Sume video does
Flow lets you build a custom voice but only for video runs that use ingredients. The Gemini API can't edit voices, and Sume video takes no voice reference.

In Google Flow you can create a custom voice, but Flow's help page says voice references can be added only to video generations that use ingredients (read 2026-10-02). Sume's video API has no voice reference at all: native audio is on for gemini-omni-flash-1.1, and there is no reference-audio field. If a specific voice matters, plan it outside the video model.
What Flow says about voices
Flow's video guide describes voice references as single-speaker audio that stays consistent across clips. You pick a preset base voice, name your voice, and describe how to change it, for example a slightly raspy voice with a New York accent (Flow Help, read 2026-10-02). The stated limitation is the one in the title: voices attach only to ingredient-based generations.
What the Gemini API says
The Gemini API page for Omni Flash lists audio uploads as unsupported in the current version and voice editing as unsupported (Gemini API docs, read 2026-10-02). So the Flow voice workflow is a Flow product feature, not something the API page documents.
| Surface | Voice reference or edit | Source |
|---|---|---|
| Google Flow | Custom voice from a preset base, ingredient runs only | Flow Help |
| Gemini API, Omni Flash | Audio upload and voice editing unsupported | Gemini API docs |
Sume gemini-omni-flash-1.1 | None: native audio always on, no reference audio field | Sume docs |
What Sume does instead
On Sume, generate_audio: false is rejected for gemini-omni-flash-1.1 and there is no reference_audio_urls (Video Router). You can steer the voice only through the prompt, such as the speaker's tone and the lines they say.
Sume's catalog notes also say video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is not a video-model clip with narration laid on top (Models overview). For a consistent on-camera voice across many clips, the honest answer is that Flow's voice feature has no Sume video-model equivalent.
When to pick which
Choose Flow if a recurring character with one voice across a series is the goal and you are comfortable inside the Flow app. Choose Sume if you need the clip from an API with durable job ids and a webhook, and the voice only has to sound right once per clip.
Sources
Related posts
More in Comparisons
- Flow Tools vs Sume Formats: a saved creative tool you can call by API
Google Flow Tools lets users build custom creative tools from plain language and share them. A Sume Format is a saved recipe you call over HTTP.
- FLUX 3 Image reference size: 256 px to 16 MP, and Sume URL rules
BFL says each FLUX 3 Image reference is 256x256 px to 16 MP, 1 to 10 images. Sume lists FLUX.2 ids; its reference rules are public HTTPS and a catalog count.
- Free speech-to-text API credits compared, October 2026
Amazon, Azure, Rev AI, Gladia, AssemblyAI, Speechmatics and Deepgram free allowances side by side, and what the same audio costs on Sume STT at $0.01 a minute.
- Free text-to-speech API allowances compared, October 2026
Azure's 0.5M free characters, Speechmatics' first 1M free and Deepgram's $200 credit side by side, plus Sume TTS limits: 20,000 characters, 1,200 seconds.
Written by Sume