Which AI video models on Sume make their own audio track
Gemini Omni Flash 1.1 always generates audio and rejects generate_audio false; MiniMax H3 Max has native stereo audio; Seedance 2.0 reports generate_audio true.

Three Sume video rows are documented as producing audio: Gemini Omni Flash 1.1, where native synced audio is always on and generate_audio: false is rejected; MiniMax H3 Max, which has native stereo audio; and Seedance 2.0, whose model entry in the docs shows generate_audio: true. For any other row, read the generate_audio field from GET /v1/videos/models before you rely on sound.
Google's pricing page quotes Veo 3.1 prices as video-with-audio prices, so audio-inclusive pricing is the norm at the top of the market; on Sume the question is only whether the row you picked makes any.
Documented audio behavior
From the Sume video docs, read 2026-10-09.
| Row | Audio | How to control it | Input audio |
|---|---|---|---|
| Gemini Omni Flash 1.1 | native synced audio, always on | generate_audio false is rejected | no audio input; no reference_audio_urls |
| MiniMax H3 Max | native stereo audio | check the model entry | reports audio in supported_input_references |
| Seedance 2.0 | generate_audio true | the generate_audio request field | audio_url reference supported |
What the request field does
generate_audio on /v1/videos tells the model to make an audio track or not, and the default is the audio capability of the model. If you set it to false on Omni, Sume refuses the request, because the audio is not optional there. That has a cost consequence: you cannot buy a silent Omni clip for less, and the 720p price of $0.125 per second always includes sound.
If you plan to lay your own music or voice over a clip, an always-on audio track is dead weight but not an extra charge. You can strip it later with the audio-detach tool, priced at $0.01 per job on the shared sheet, or replace the sound in a timeline render at $0.10 per output minute.
Checking before you build
Fetch the live list and filter it. The following returns the id and audio flag for each row; run it with your API key set.
curl -s "https://api.sume.com/v1/videos/models" \
-H "Authorization: Bearer $SUME_API_KEY" \
| python3 -c "import json,sys; [print(m['id'], m['generate_audio']) for m in json.load(sys.stdin)['data']]"Cost of sound
The following figures use the same list prices as the tables above.
- Omni's audio is in the $0.125 per second price, so a 10-second 720p clip with sound is $1.25.
- Seedance 2.0 at 720p is $0.378 per second, and a 10-second clip is $3.78 with or without the field set, as far as the docs state.
- Sume's price sheet lists TTS at $0.0475 per 1,000 characters if you prefer to add a voiceover separately.
Sources
Related posts
More in Developers
- Which id to send Sume support: req_, job_, arun_ or agrun_
Error bodies carry req_ ids; jobs have job_; Format and Action runs arun_; Agent Completions agrun_. Which to share with support, and what never to paste.
- Which Sume wait mode fits which job: image, video, avatar, swap
Sume's async, sync, subscribe and webhook modes differ only in how you learn the outcome. A table by job length, and why 30 seconds is a request budget.
- X API video upload: chunked only, 0.5 s minimum, 20-minute cap
X's media docs require chunked upload for all videos, set a 0.5-second minimum and a 20-minute cap for non-Premium posts. Trim a clip for DMs with video_trim.
- Wan 3.0 X-DashScope-Async header and polling vs Sume
Alibaba needs an X-DashScope-Async: enable header and suggests polling about every 15 seconds. Sume is async by default and hands you a polling_url.
Written by Sume