Shorts split audio from video: which Sume models let you drop AI sound
The Shorts editor now separates audio and video tracks. Which Sume video models can generate silent footage, which always embed audio, and the price gap.

YouTube Shorts' editor now separates the audio and video tracks, as reported by Heyorca (read 2026-10-07). For people who generate clips with AI, that moves a decision upstream: do you want the model's own sound, or do you plan to lay your own music and voice over silent footage?
On Sume the answer depends on the model. Three rows always bake audio in, three let you switch it off, and one has no audio at all.
Audio behaviour by model
| Model | Audio | Can you send generate_audio: false | Price effect |
|---|---|---|---|
| gemini-omni-flash-1.1 | always on, synced | no, the API rejects it | included in the per-second rate |
| minimax-h3 | always on, stereo | no field | included |
| minimax-h3-max | always on, stereo | no field | included |
| kling-3 | optional | yes | $0.14 per second silent, $0.21 with audio |
| wan-3.0 | optional | yes | one per-second rate by resolution |
| seedance-2.5 and the 2.0 family | optional | yes | token-priced, not split by audio in the catalog |
| grok-imagine-video-1.5 | none | not accepted | flat per second |
Which to choose for a Short you will score yourself
If you are adding music or a voiceover in the Shorts editor, buy silent footage. Kling 3 is the only row where silence is cheaper in the catalog: a 15-second silent clip is $2.10 against $3.15 with audio. For Wan 3.0 and Seedance the catalog does not price audio separately, so turning it off saves nothing, but it does keep a stray AI voice out of your mix.
If you want the model's sound, use Omni Flash. At 720p a 10-second clip is $1.25 with synced audio, and the audio cannot be separated before you download. In the Shorts editor you can then drop that track and keep the picture, because the new editor handles the two tracks apart.
A silent request
Send the flag explicitly. A model that does not accept it returns an error, which tells you the row has fixed audio.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"kling-3","prompt":"Barista pours latte art, overhead, no speech","duration":8,"resolution":"1080p","aspect_ratio":"9:16","generate_audio":false}'Sources
Related posts
More in Use cases
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
- Talking avatar for PowerPoint presentations, slide by slide
Make a talking avatar presenter for PowerPoint: one Sume clip per slide, up to 60 seconds each, in 16:9 or 4:3 to match the slide, inserted as MP4.
- AI avatar for YouTube videos: Shorts and long-form
Use an AI avatar in YouTube videos: a 9:16 talking video of up to 60 seconds for a Short, or 16:9 segments joined into one long-form video.
Written by Sume