Native audio in stitched AI clips: which Sume models force it
Gemini Omni Flash 1.1 always makes audio, H3 Max adds stereo, Seedance 2 can. A stitched Timeline render takes its audio from the spine, so plan the sound.

On Sume, Gemini Omni Flash 1.1 always generates audio (the API rejects generate_audio: false), MiniMax H3 Max returns native stereo audio, and Seedance 2 lists audio generation as a capability. When you stitch clips in Timeline 1.0, the render takes its sound from the audio spine you declare, so decide the soundtrack before you join.
This matters for a multi-shot cut like the 16-keyframe sequences Luma markets for Ray 3.2, which Sume does not run: sixteen clips each with their own sound will not sit together.
Which models make sound
The Video generation docs list generate_audio per model in GET /v1/videos/models; the default is the model's audio capability. The Video Router docs add that Omni Flash 1.1 has native synced audio always on, with no reference_audio_urls, and that minimax-h3-max has native stereo audio.
| Model | Audio behavior | Reference audio in |
|---|---|---|
| gemini-omni-flash-1.1 | always on; generate_audio false is rejected | no |
| minimax-h3-max | native stereo audio | yes (audio refs reported) |
| seedance-2 | generate_audio true in the catalog entry | yes |
| Seedance 2.x, Wan 3.0, MiniMax H3 | accept audio and video references | yes |
| higgsfield-genjutsu, h3-max-recast | video references only | no |
| Kling motion control | keep_original_sound default true; false is silent | n/a |
What Timeline does with it
Timeline 1.0 requires an audio spine: a url, up to 20 parts, or mode: silence. The compose docs say that a mute video only warns (compose_video_has_no_audio), and that the Timeline 1.0 spine supplies the audio at assemble time. Treat each clip's own sound as something the spine replaces, and listen to the result rather than assuming the clip audio survives the join.
A soundtrack bed can sit under the spine, with gain_db, loop, fade_out_seconds up to 10 and duck_db of 0 to 20. Ducking needs a real spine, not silence.
Three workable plans
- Voiceover plan: generate a narration, use it as
audio.url, add a musicsoundtrackwithduck_dbso the bed drops under the voice. - Silent plan: use
audio.mode: silence, then add sound in your editor. Quiet clips avoid a clash between models' built-in audio. - Single-model plan: keep one audio-capable model across the chain so the generated sound has one character, and still verify the result.
Cost angle
Audio is part of the provider list for models where it is on, and Sume bills that list times 1.25 per output second. For Omni Flash 1.1 you cannot turn audio off to save money, because the API rejects it. For a silent edit that you will score yourself, choose an id where audio is optional, or use Kling's keep_original_sound: false.
The Timeline render itself is $0.10 per ceil(output minute) and runs worker ffmpeg only.
Check the result
After the render, run video inspect on the output: its probe includes has_audio, so you can tell whether the file has an audio track at all. Then listen at the seams. If you plan a detached stem, POST /v1/audio-detach returns a WAV or an MP3 of the track.
If you need the original clip sound under a voiceover, detach it from each clip and place the pieces with audio.parts[] (up to 20 gapless slices) rather than hoping the video audio carries through.
Why this matters for a pro pipeline
A finishing editor usually wants control of the mix: dialogue, music and effects on separate tracks. Clips with baked-in generated sound cut against that. Where the model lets you, choose silent generation, then build the mix from a voiceover spine and a ducked soundtrack bed. Where it does not, detach the audio and decide per clip whether to keep it.
Before you build
Before you build, read the linked Sume docs page for the exact request fields, limits and prices, because those pages are the source of truth and can change. Run one short, cheap test with your own material first, check the output in a player and in your editor, and only then scale to the full shot list. Keep every job id and file you approve, so a later change never forces you to regenerate work that was already signed off. Note that this post describes Sume's catalog and tools; Sume does not run Luma Ray 3.2, and nothing here claims HDR or EXR output.
Sources
Related posts
More in Developers
- Next.js route handlers to replace a Sora call: submit + webhook
Two App Router route handlers: one posts a video job to Sume with an idempotency key and callback_url, one verifies the signed job.completed webhook.
- Next.js Pages Router webhook for a Sume video job: bodyParser false
A Pages Router API route that reads the raw body, verifies sume-v1, acks fast and dedupes on job_id for a 30 s render. Refuses an empty secret.
- No GET /v1/audio-detach/:id: poll the job instead
Audio detach and timeline audio have no GET resource route. Read /v1/jobs/:id/status and /result. Video captions and avatar video do have resource GETs.
- Node 22 batch runner for a prompt file: four lanes on Sume
Read one prompt per line, render each on Sume POST /v1/videos with four concurrent lanes and a stable idempotency key per line, and save clip-N.mp4. Node 22.
Written by Sume