Which AI video models on Sume make their own audio track

Gemini Omni Flash 1.1 always generates audio and rejects generate_audio false; MiniMax H3 Max has native stereo audio; Seedance 2.0 reports generate_audio true.

5 min readSume
All posts

Three Sume video rows are documented as producing audio: Gemini Omni Flash 1.1, where native synced audio is always on and generate_audio: false is rejected; MiniMax H3 Max, which has native stereo audio; and Seedance 2.0, whose model entry in the docs shows generate_audio: true. For any other row, read the generate_audio field from GET /v1/videos/models before you rely on sound.

Google's pricing page quotes Veo 3.1 prices as video-with-audio prices, so audio-inclusive pricing is the norm at the top of the market; on Sume the question is only whether the row you picked makes any.

Documented audio behavior

From the Sume video docs, read 2026-10-09.

Audio capability of Sume video rows, as of 2026-10-09
RowAudioHow to control itInput audio
Gemini Omni Flash 1.1native synced audio, always ongenerate_audio false is rejectedno audio input; no reference_audio_urls
MiniMax H3 Maxnative stereo audiocheck the model entryreports audio in supported_input_references
Seedance 2.0generate_audio truethe generate_audio request fieldaudio_url reference supported

What the request field does

generate_audio on /v1/videos tells the model to make an audio track or not, and the default is the audio capability of the model. If you set it to false on Omni, Sume refuses the request, because the audio is not optional there. That has a cost consequence: you cannot buy a silent Omni clip for less, and the 720p price of $0.125 per second always includes sound.

If you plan to lay your own music or voice over a clip, an always-on audio track is dead weight but not an extra charge. You can strip it later with the audio-detach tool, priced at $0.01 per job on the shared sheet, or replace the sound in a timeline render at $0.10 per output minute.

Checking before you build

Fetch the live list and filter it. The following returns the id and audio flag for each row; run it with your API key set.

curl -s "https://api.sume.com/v1/videos/models" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  | python3 -c "import json,sys; [print(m['id'], m['generate_audio']) for m in json.load(sys.stdin)['data']]"

Cost of sound

The following figures use the same list prices as the tables above.

  • Omni's audio is in the $0.125 per second price, so a 10-second 720p clip with sound is $1.25.
  • Seedance 2.0 at 720p is $0.378 per second, and a 10-second clip is $3.78 with or without the field set, as far as the docs state.
  • Sume's price sheet lists TTS at $0.0475 per 1,000 characters if you prefer to add a voiceover separately.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume