Reference audio for Wan 3.0 and MiniMax H3: WAV, MP3, 15 MB
Both Wan 3.0 and MiniMax H3 take WAV or MP3 reference audio up to 15 MB and 15 seconds in total. Where the clip counts differ and what Sume says.

Both vendors accept WAV and MP3 reference audio with a 15 MB per-file cap and 15 seconds of audio in total, but they count clips differently: Alibaba allows up to 5 clips for Wan 3.0, MiniMax up to 3 for H3 and H3 Max. Sume documents that Wan 3.0, MiniMax H3 and MiniMax H3 Max honor audio and video references, and that the legacy Video Router Auto shape allows 1 to 3 reference audio URLs.
Figures are from Alibaba's Wan3.0 Video Generation API Reference and MiniMax's Create Video Generation Task, both read 2026-10-02. Sume's are from Video generation and Video Router.
How do the two vendors compare?
The shared parts are format, size and total length. The differences are the clip count and MiniMax's per-clip minimum of 2 seconds. Alibaba's page gives no per-clip minimum for audio in the text I fetched, so do not assume one.
| Item | Wan 3.0 (Alibaba) | MiniMax H3 / H3 Max |
|---|---|---|
| Formats | wav, mp3 | WAV, MP3 |
| Max file size | 15 MB | 15 MB |
| Clips | up to 5 | up to 3 |
| Total length | 15 seconds or less | 15 seconds or less |
| Per clip | not stated in text read | 2 to 15 seconds |
| Role in request | reference_audio | reference_audio |
What does Sume say about audio references?
Two things. First, audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1 and Higgsfield Genjutsu accept video references but not audio. Second, on the legacy Video 1.0 and Auto shape, reference_audio_urls takes 1 to 3 URLs and requires at least one reference image or video alongside it. MiniMax makes a similar point: audio is a reference type, used with other references, not a standalone input.
The Sume pages I read do not publish the same 5-clip figure for Wan, so with Sume the conservative cap is 3 audio clips, under both vendors' limits.
How do I submit one safely?
Use public HTTPS URLs, one image or video reference next to the audio, and keep total audio at 15 seconds. This example uses the Video Router's flat fields.
import os, requests
r = requests.post(
"https://api.sume.com/v1/video-router/generate",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "audio-ref-001"},
json={"model": "minimax-h3", "duration": 10, "resolution": "768p",
"prompt": "The host from the photo speaks to camera in a warm voice",
"reference_image_urls": ["https://example.com/host.png"],
"reference_audio_urls": ["https://example.com/voice.mp3"],
"mode": "async"},
)
print(r.status_code, r.json())What should I verify before relying on it?
The model-by-model audio list names which ids take audio on Sume.
- Read
supported_input_referencesfor the model id atGET /v1/videos/models; only listed types are accepted. - Confirm the audio is WAV or MP3, 15 MB or smaller and 15 seconds total.
- Do not use a reference audio when you want Sume to generate speech from the prompt; that is a different input.
- Remember MiniMax's own rule that H3 reference audio is a reference, not a lip-sync track.
Sources
Related posts
More in Media tools
- Wan 3.0 reference images: 20 MB, 240 to 8,000 px, 8:1 cap
Alibaba's Wan 3.0 reference images must be 20 MB or less, 240 to 8,000 px per side, ratio 8:1 at most. A check to run before wan-3.0 on Sume.
- Which AI video models take an end frame in Sume's video picker?
Wan 3.0, MiniMax H3, H3 Max and Auto take an end frame in Sume's Videos panel; Kling 3.0 and Grok Imagine do not. How to check the API list too.
- YouTube caption file for a Short: with timing or without timing?
YouTube's Upload file option asks for With timing or Without timing. Which to pick for a Short, and how Sume's transcript segments and burned-in captions fit.
- YouTube Shorts max 1080p: downscale a 4K vertical clip with Sume
YouTube's Shorts help says uploads have a maximum resolution of 1080p. Downscale a 2160x3840 clip to 1080x1920 with video-trim's exact-mode output conform.
Written by Sume