Captions on AI video with native sound: what Sume actually transcribes

Gemini Omni Flash 1.1 always generates synced audio on Sume. The captions endpoint transcribes speech only, so sound effects need authored cues.

4 min readSume
All posts

Run the generated clip through the captions endpoint: where the model produced speech, it is transcribed and styled; where it produced only music or sound effects, there is nothing to transcribe. If the whole clip has no speech you get caption_no_speech, and you supply cues yourself.

Gemini Omni Flash 1.1 always generates synced audio on Sume, so every clip has a soundtrack, but not every soundtrack has words.

What the video docs say

The Video Router docs state that Gemini Omni Flash 1.1 has native synced audio that is always on: generate_audio false is rejected and reference_audio_urls is not accepted. Seedance 2.x, Wan 3.0 and the MiniMax H3 models accept audio references instead. Video Router billing is list price times 1.25, as the docs state.

That means for Omni you cannot request a silent clip to caption later; you can only ignore or replace the audio afterwards.

What captions will and will not do

The captions endpoint charges $0.20 for videos up to 60 seconds. It transcribes speech with a language hint that never changes the style or font. A clip with no speech returns caption_no_speech, in which case pass cues or segments. Korean text on a Latin style returns caption_hangul_text_latin_style, so match the style to the script of the words.

  • Speech in the clip: let it transcribe, review, then render.
  • No speech: author cues for the lines or labels you want.
  • Mixed: build one cue list from your script and send it.
  • Long clips: stay at or under 60 seconds.

Check the audio first

Before paying for captions, listen to the clip. Generated voices can mumble or drift from the prompted line, and a transcript will faithfully capture that. If the spoken line must match your script exactly, pass script_text to align the script to the audio; script_alignment_mismatch warns you when the model said something else.

What each source of audio gives the captions step (read 2026-10-03)
Audio in the clipCaptions result
Clear speechTranscribed and styled
Music and effects onlycaption_no_speech; use cues
Speech with script_textAligned; mismatch error if different

Pairing with a music bed

If you add a soundtrack on top of the native audio, duck it under the voice at render time. The duck post shows how to detach the native track first. Do the captions before the mix; the mix does not change the words.

Related posts

More in Use cases

All Use cases posts

Written by Sume