AI video API with audio reference input: which Sume models accept it

Seedance (all four), Wan 3.0, MiniMax H3 and H3 Max accept audio references on Sume. Kling, Gemini Omni, Grok Imagine and the motion models do not.

5 min readSume
All posts

Four model families on Sume accept reference audio: the Seedance 2.x family (2.5, 2.0, Fast, Mini), Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1, Kling Video v3 Pro, Grok Imagine Video 1.5, Genjutsu and H3 Max Recast do not. Sume's docs say the Seedance 2.x models, Wan 3.0 and the MiniMax models accept audio and video references, while Gemini Omni Flash 1.1, Genjutsu and Recast accept video references but not audio. Read as of 2026-10-08.

A reference audio clip is an input that guides the generated sound, so it matters for lip-sync style shots, music-driven edits and voice-led ads.

Audio reference limits by model (Sume docs and catalog, as of 2026-10-08)
ModelReference audioLimit
wan-3.0yesup to 5 clips, 15 s total
minimax-h3, minimax-h3-maxyesup to 3 clips, 2 to 15 s each, 15 s combined; audio cannot be the only reference
seedance-2.5 and 2.0 familyyes1 to 3 clips (Video 1.0 shape), with at least one reference image or video
gemini-omni-flash-1.1novideo and image references only
kling-3nono reference inputs at all
grok-imagine-video-1.5nofirst-frame image only
higgsfield-genjutsu, h3-max-recastnoimage and video only

Sending the audio

On the Video 1.0 shape, send reference_audio_urls (1 to 3 URLs). You must also send at least one reference image or video: audio alone is not a valid reference. On /v1/videos, add an audio_url entry to input_references for models whose supported_input_references lists it.

curl -X POST https://api.sume.com/v1/video-1.0/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"The singer performs to camera","reference_image_urls":["https://example.com/singer.png"],"reference_audio_urls":["https://example.com/take.mp3"],"resolution":"720p","duration":8}'

Price effects

Audio references add no line item on Wan 3.0 or MiniMax H3 Max, which bill output seconds only: 8 s of Wan 3.0 at 720p is $1.00 and 8 s of H3 Max at 768p is $0.80. On Seedance, a reference video triggers the 15-second assumption and the 0.6 multiplier, but the pricing code looks at reference video fields, not audio. For the Video 1.0 route, the Auto family always generates audio and the docs say not to send generate_audio.

A worked example

Say you want an 8-second 9:16 clip where a presenter speaks along to a 10-second voice take. On Wan 3.0 at 720p the clip costs 8 x $0.125 = $1.00. On MiniMax H3 at 768p it costs 8 x $0.075 = $0.60, plus $0.10 for each reference image beyond the fifth. Both models accept a 10-second audio file, since the limit on each is 15 seconds combined. Seedance 2.0 Fast would accept the audio too, but the price depends on whether you also send a reference video.

Keep audio references clean and short, and always attach at least one image or video reference alongside, since audio cannot be the only reference on MiniMax or the Seedance Video 1.0 shape.

Which to pick

  • Long audio: Wan 3.0 takes up to 15 seconds across five clips.
  • Strict 2 to 15 second clips: MiniMax, where each clip must be at least 2 seconds.
  • Need no audio input: skip the field and let the model generate synced sound.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume