Sound effects for video AI: write the sounds into the clip prompt

With no sound-effects endpoint, video effects come from models that generate audio with the clip. Check the generate_audio flag, then describe each sound.

4 min readSume
All posts

For sound effects on an AI video, ask the video model to make them: on Sume, models whose catalog entry has generate_audio can produce an audio track with the clip, and you describe each sound in the prompt. There is no separate sound-effects endpoint to call afterwards, so a model without audio gives you a silent clip that you score in Timeline 1.0.

The facts are from the Video generation and Timeline 1.0 docs, read 2026-09-29.

Which video models can make sound?

The catalog flag is generate_audio, described as whether the model can generate an audio track. The docs name these examples:

From the Sume Video generation docs, read 2026-09-29.
ModelWhat the docs say about audio
seedance-2Text, image and reference to video with optional audio; generate_audio is true
minimax-h3-maxNative stereo audio
gemini-omni-flash-1.1Native synced audio

How do I write the sounds?

Say what is heard next to what is seen, in the order it happens, and set generate_audio to true. The request field defaults to the model's audio capability. Read the returned clip with the sound on before you commit.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sfx-bar-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "A glass slides along a wooden bar and stops at a hand. Ice clinks. Low room murmur under it.",
    "duration": 5,
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "generate_audio": true
  }'

What if the clip is silent or the sound is wrong?

Score it after the fact. Timeline 1.0 takes a soundtrack bed with url, gain_db, loop, fade_out_seconds and duck_db, and every file must be a Sume-hosted media.sume.com file, so import outside audio first with POST /v1/media-imports. duck_db needs a real audio spine, not silence.

Can I send my own sound as a reference?

Some models accept an audio reference. Sume's docs say only models whose supported_input_references lists a type accept that type, and that audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1 accepts video references but not audio.

List GET /v1/videos/models to see each model's flags before you write a request, since the catalog changes.

How do I keep the sound and picture matched?

Generate several short clips instead of one long one, and describe one action and its sound per clip. Review each with sound on, then join the approved clips in Timeline 1.0. Shorter clips make a wrong sound cheaper to redo, and you keep the good takes.

Sume's docs make no promise about how closely generated sound follows a prompt, so treat each clip as a draft until you have listened to it.

What should I write in the prompt?

Name the source of each sound, not just the sound. “Ice clinks against glass” gives the model an object and an action, while “clinking” gives it a noise with no cause. Keep the list short: two or three sounds per clip, one of them dominant.

The catalog and the docs describe capability, not quality, so the clip itself is the test.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume