Sound effects for video AI: write the sounds into the clip prompt
With no sound-effects endpoint, video effects come from models that generate audio with the clip. Check the generate_audio flag, then describe each sound.

For sound effects on an AI video, ask the video model to make them: on Sume, models whose catalog entry has generate_audio can produce an audio track with the clip, and you describe each sound in the prompt. There is no separate sound-effects endpoint to call afterwards, so a model without audio gives you a silent clip that you score in Timeline 1.0.
The facts are from the Video generation and Timeline 1.0 docs, read 2026-09-29.
Which video models can make sound?
The catalog flag is generate_audio, described as whether the model can generate an audio track. The docs name these examples:
| Model | What the docs say about audio |
|---|---|
seedance-2 | Text, image and reference to video with optional audio; generate_audio is true |
minimax-h3-max | Native stereo audio |
gemini-omni-flash-1.1 | Native synced audio |
How do I write the sounds?
Say what is heard next to what is seen, in the order it happens, and set generate_audio to true. The request field defaults to the model's audio capability. Read the returned clip with the sound on before you commit.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sfx-bar-001" \
-d '{
"model": "seedance-2",
"prompt": "A glass slides along a wooden bar and stops at a hand. Ice clinks. Low room murmur under it.",
"duration": 5,
"resolution": "720p",
"aspect_ratio": "16:9",
"generate_audio": true
}'What if the clip is silent or the sound is wrong?
Score it after the fact. Timeline 1.0 takes a soundtrack bed with url, gain_db, loop, fade_out_seconds and duck_db, and every file must be a Sume-hosted media.sume.com file, so import outside audio first with POST /v1/media-imports. duck_db needs a real audio spine, not silence.
Can I send my own sound as a reference?
Some models accept an audio reference. Sume's docs say only models whose supported_input_references lists a type accept that type, and that audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1 accepts video references but not audio.
List GET /v1/videos/models to see each model's flags before you write a request, since the catalog changes.
How do I keep the sound and picture matched?
Generate several short clips instead of one long one, and describe one action and its sound per clip. Review each with sound on, then join the approved clips in Timeline 1.0. Shorter clips make a wrong sound cheaper to redo, and you keep the good takes.
Sume's docs make no promise about how closely generated sound follows a prompt, so treat each clip as a draft until you have listened to it.
What should I write in the prompt?
Name the source of each sound, not just the sound. “Ice clinks against glass” gives the model an object and an action, while “clinking” gives it a noise with no cause. Keep the list short: two or three sounds per clip, one of them dominant.
The catalog and the docs describe capability, not quality, so the clip itself is the test.
Sources
Related posts
More in Media tools
- 24fps vs 30fps AI video: which output.fps to set on Sume
Should an AI video be 24 or 30 fps? On Sume's Timeline, leave output.fps unset to follow your sources; set 24, 25, 30 or 60 only when a delivery spec needs it.
- Reframe an AI video to 9:16: regenerate, crop or fit?
Luma lists enhanced reframe for Ray3.2. On Sume, a 9:16 frame comes from asking for it, cropping with video filter, or fitting in Timeline.
- Amazon online video ad specs: OLV size, length, bitrate
Amazon online video (OLV) ads run 6–120 s in 16:9, at least 1920×1080 and 4 Mbps, with 192 kbps AAC on 2+ channels and up to 500 MB site-served.
- App Store screenshot size (1290×2796) and AI images
Apple lists 1290×2796 for 6.9-inch iPhone screenshots. Sume's ChatGPT Image 2.5 cannot output it exactly, so generate the same shape larger and resize.
Written by Sume