Vidu Q4 takes 3 audio references; which Sume video models take one?

Sume does not list Vidu Q4 Preview. Seedance 2.x, Wan 3.0 and MiniMax H3 accept reference audio; Gemini Omni Flash 1.1 does not. Prices for 15 seconds.

5 min readSume
All posts

If you want a branded sound to carry across several ad cuts, you need a video model that accepts a reference track. Vidu Q4 Preview, launched 2026-10-07, takes up to 3 audio references but is not in Sume's video catalog. On Sume, Seedance 2.5, Seedance 2.0 and its Fast and Mini variants, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept reference audio; Gemini Omni Flash 1.1 does not (read 2026-10-08).

What each side documents

Vidu's release says Q4 Preview takes up to 15 image references and up to 3 audio references, with output from 540p to 4K and pricing from $0.014 per second. Sume's docs say a model accepts a reference type only if its supported_input_references includes it. The Seedance 2.x models, Wan 3.0 and the MiniMax H3 pair accept audio and video references. Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio.

Reference audio, as of 2026-10-08
ModelReference audioDocumented limit
Vidu Q4 Preview (Vidu, not on Sume)YesUp to 3
seedance-2.5, seedance-2, -fast, -miniYesRead the catalog entry
wan-3.0YesUp to 5 clips, 15 s total
minimax-h3, minimax-h3-maxYesRead the catalog entry
gemini-omni-flash-1.1NoNative synced audio always on

Price of a 15-second clip with a track

A reference track does not change the per-second price in the Sume catalog data. These are billable prices per 15-second clip at 720p, for models that accept a reference track and that have a 15-second length.

15 seconds at 720p as of 2026-10-08
ModelBillable per clip
wan-3.0$1.88
seedance-2-mini (9:16)$2.84
seedance-2.5 (9:16)$8.67

Using the track

Pass the track as an audio reference on POST /v1/videos, with a public HTTPS URL. Treat the output as inspired by the reference, not a guaranteed lip or beat match, and listen before you ship. If the sound must be exact, generate the clip with the model's own audio and lay your track on Timeline 1.0 with soundtrack, which lets you set gain_db, loop and duck_db.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-ref-01" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Product spins on a turntable to the beat",
    "resolution": "720p",
    "duration": 15,
    "input_references": [
      { "type": "audio_url", "audio_url": { "url": "https://media.sume.com/artifacts/artf_demo/brand-sting.mp3" } }
    ]
  }'

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume