Captions on AI video with native sound: what Sume actually transcribes
Gemini Omni Flash 1.1 always generates synced audio on Sume. The captions endpoint transcribes speech only, so sound effects need authored cues.

Run the generated clip through the captions endpoint: where the model produced speech, it is transcribed and styled; where it produced only music or sound effects, there is nothing to transcribe. If the whole clip has no speech you get caption_no_speech, and you supply cues yourself.
Gemini Omni Flash 1.1 always generates synced audio on Sume, so every clip has a soundtrack, but not every soundtrack has words.
What the video docs say
The Video Router docs state that Gemini Omni Flash 1.1 has native synced audio that is always on: generate_audio false is rejected and reference_audio_urls is not accepted. Seedance 2.x, Wan 3.0 and the MiniMax H3 models accept audio references instead. Video Router billing is list price times 1.25, as the docs state.
That means for Omni you cannot request a silent clip to caption later; you can only ignore or replace the audio afterwards.
What captions will and will not do
The captions endpoint charges $0.20 for videos up to 60 seconds. It transcribes speech with a language hint that never changes the style or font. A clip with no speech returns caption_no_speech, in which case pass cues or segments. Korean text on a Latin style returns caption_hangul_text_latin_style, so match the style to the script of the words.
- Speech in the clip: let it transcribe, review, then render.
- No speech: author cues for the lines or labels you want.
- Mixed: build one cue list from your script and send it.
- Long clips: stay at or under 60 seconds.
Check the audio first
Before paying for captions, listen to the clip. Generated voices can mumble or drift from the prompted line, and a transcript will faithfully capture that. If the spoken line must match your script exactly, pass script_text to align the script to the audio; script_alignment_mismatch warns you when the model said something else.
| Audio in the clip | Captions result |
|---|---|
| Clear speech | Transcribed and styled |
| Music and effects only | caption_no_speech; use cues |
| Speech with script_text | Aligned; mismatch error if different |
Pairing with a music bed
If you add a soundtrack on top of the native audio, duck it under the voice at render time. The duck post shows how to detach the native track first. Do the captions before the mix; the mix does not change the words.
Related posts
More in Use cases
- Should captions censor profanity? Section 508 and YouTube's [ __ ]
Section 508 says caption audible profanity exactly; YouTube's auto captions swap flagged words for [ __ ] by default. How to burn the wording you intend.
- Unreadable captions on bright Reels: dim 0.45, check for free
Burned-in captions vanish on bright footage. Sume video filter dim 0.45 darkens luma; the /video-filter/check endpoint validates the program at no charge.
- Captions covering Shorts stickers: move the line with anchor_ratio
Burned-in captions can sit where a Shorts sticker will land. Sume's caption design.placement.anchor_ratio moves the line, from 0.05 to 0.95 of frame height.
- Change a car color in a video with Gemini Omni Flash edit
Try recoloring a car in a clip with Gemini Omni Flash 1.1 on Sume: object swap prompt, reflection check, 10 second limit, request and video inspect QA.
Written by Sume