Shorts split audio from video: which Sume models let you drop AI sound

The Shorts editor now separates audio and video tracks. Which Sume video models can generate silent footage, which always embed audio, and the price gap.

5 min readSume
All posts

YouTube Shorts' editor now separates the audio and video tracks, as reported by Heyorca (read 2026-10-07). For people who generate clips with AI, that moves a decision upstream: do you want the model's own sound, or do you plan to lay your own music and voice over silent footage?

On Sume the answer depends on the model. Three rows always bake audio in, three let you switch it off, and one has no audio at all.

Audio behaviour by model

From the Sume model cards and catalog, read 2026-10-07.
ModelAudioCan you send generate_audio: falsePrice effect
gemini-omni-flash-1.1always on, syncedno, the API rejects itincluded in the per-second rate
minimax-h3always on, stereono fieldincluded
minimax-h3-maxalways on, stereono fieldincluded
kling-3optionalyes$0.14 per second silent, $0.21 with audio
wan-3.0optionalyesone per-second rate by resolution
seedance-2.5 and the 2.0 familyoptionalyestoken-priced, not split by audio in the catalog
grok-imagine-video-1.5nonenot acceptedflat per second

Which to choose for a Short you will score yourself

If you are adding music or a voiceover in the Shorts editor, buy silent footage. Kling 3 is the only row where silence is cheaper in the catalog: a 15-second silent clip is $2.10 against $3.15 with audio. For Wan 3.0 and Seedance the catalog does not price audio separately, so turning it off saves nothing, but it does keep a stray AI voice out of your mix.

If you want the model's sound, use Omni Flash. At 720p a 10-second clip is $1.25 with synced audio, and the audio cannot be separated before you download. In the Shorts editor you can then drop that track and keep the picture, because the new editor handles the two tracks apart.

A silent request

Send the flag explicitly. A model that does not accept it returns an error, which tells you the row has fixed audio.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"kling-3","prompt":"Barista pours latte art, overhead, no speech","duration":8,"resolution":"1080p","aspect_ratio":"9:16","generate_audio":false}'

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume