Vidu Q4 Preview takes 3 audio references: which Sume rows take audio

Vidu Q4 Preview accepts up to 3 mp3 voice references of 3 to 12 s. On Sume, Seedance 2.x, Wan 3.0 and the MiniMax H3 rows take audio references; Omni does not.

4 min readSume
All posts

Vidu Q4 Preview accepts up to three audio reference clips so a character keeps one voice across shots. Vidu's reference-to-video page lists 0 to 3 audio files in mp3 format, 3 to 12 seconds each and up to 50 MB per file. If you want the same idea on Sume, five catalog rows list audio among their input references: the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Gemini Omni Flash 1.1, Kling 3, Genjutsu and H3 Max Recast do not take audio references.

What each side documents

The launch release says up to three reference audio clips guide voice consistency. Sume's video guide says the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references, while Gemini Omni Flash 1.1, higgsfield-genjutsu and h3-max-recast accept video references but not audio.

Audio-reference support, read 2026-10-08
ModelAudio referencesDocumented limit
Vidu Q4 Preview (Vidu API)Yes0 to 3 mp3 files, 3 to 12 s each
wan-3.0 (Sume)Yesup to 5 files, up to 15 s combined
minimax-h3 / minimax-h3-max (Sume)Yesup to 3 files, 2 to 15 s each, 15 s combined; audio cannot be the only reference
seedance-2.5 and 2.x (Sume)Yesread the row's capabilities in the catalog
gemini-omni-flash-1.1 (Sume)Nonative synced audio is always generated
kling-3 (Sume)Nono reference inputs

Send an audio reference on Sume

Audio goes in input_references as an audio_url item, next to at least one image or video, because on the H3 rows audio alone is rejected as the only reference. Check that your target row lists audio_url before you send it.

Steps:

  • Export the voice sample as a short clip that is public over HTTPS.
  • Add an image_url entry for the character.
  • Add an audio_url entry for the voice.
  • Submit to POST /v1/videos with model set to wan-3.0 or minimax-h3-max.
{
  "model": "minimax-h3-max",
  "prompt": "She greets the camera and explains the product",
  "resolution": "768p",
  "duration": 8,
  "input_references": [
    {"type": "image_url", "image_url": {"url": "https://example.com/host.png"}},
    {"type": "audio_url", "audio_url": {"url": "https://example.com/host-voice.mp3"}}
  ]
}

Picking a Sume row for a voiced character

If the clip is a talking character, start with the rows that accept audio: Wan 3.0 (480p, 720p or 1080p, 2 to 30 seconds, $0.0625, $0.125 or $0.25 a second) or MiniMax H3 Max (480p, 768p or 1080p, 5 to 15 seconds, $0.0625, $0.10 or $0.20 a second). Both produce native audio, and the H3 rows have no audio on/off toggle, because stereo sound is always produced.

If you only need a person to speak a script on camera, the Avatar talking-video route is a better fit than a reference clip, since it takes a script and a ready avatar and produces a 4 to 60 second video. That route is documented separately from the video router.

What Sume does not do

Sume does not claim that an audio reference gives the same voice match as Vidu's feature, because the two models differ and Sume documents no voice-consistency guarantee. It also has no Vidu row. If you need a locked voice and do not want to depend on a reference, the Sume TTS router produces a voice track separately, but then you join sound and picture yourself.

Vidu notes that final feature availability can vary by plan and region.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume