Reference audio for Wan 3.0 and MiniMax H3: WAV, MP3, 15 MB

Both Wan 3.0 and MiniMax H3 take WAV or MP3 reference audio up to 15 MB and 15 seconds in total. Where the clip counts differ and what Sume says.

5 min readSume
All posts

Both vendors accept WAV and MP3 reference audio with a 15 MB per-file cap and 15 seconds of audio in total, but they count clips differently: Alibaba allows up to 5 clips for Wan 3.0, MiniMax up to 3 for H3 and H3 Max. Sume documents that Wan 3.0, MiniMax H3 and MiniMax H3 Max honor audio and video references, and that the legacy Video Router Auto shape allows 1 to 3 reference audio URLs.

Figures are from Alibaba's Wan3.0 Video Generation API Reference and MiniMax's Create Video Generation Task, both read 2026-10-02. Sume's are from Video generation and Video Router.

How do the two vendors compare?

The shared parts are format, size and total length. The differences are the clip count and MiniMax's per-clip minimum of 2 seconds. Alibaba's page gives no per-clip minimum for audio in the text I fetched, so do not assume one.

Reference audio limits (read 2026-10-02)
ItemWan 3.0 (Alibaba)MiniMax H3 / H3 Max
Formatswav, mp3WAV, MP3
Max file size15 MB15 MB
Clipsup to 5up to 3
Total length15 seconds or less15 seconds or less
Per clipnot stated in text read2 to 15 seconds
Role in requestreference_audioreference_audio

What does Sume say about audio references?

Two things. First, audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, while Gemini Omni Flash 1.1 and Higgsfield Genjutsu accept video references but not audio. Second, on the legacy Video 1.0 and Auto shape, reference_audio_urls takes 1 to 3 URLs and requires at least one reference image or video alongside it. MiniMax makes a similar point: audio is a reference type, used with other references, not a standalone input.

The Sume pages I read do not publish the same 5-clip figure for Wan, so with Sume the conservative cap is 3 audio clips, under both vendors' limits.

How do I submit one safely?

Use public HTTPS URLs, one image or video reference next to the audio, and keep total audio at 15 seconds. This example uses the Video Router's flat fields.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/video-router/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "audio-ref-001"},
    json={"model": "minimax-h3", "duration": 10, "resolution": "768p",
          "prompt": "The host from the photo speaks to camera in a warm voice",
          "reference_image_urls": ["https://example.com/host.png"],
          "reference_audio_urls": ["https://example.com/voice.mp3"],
          "mode": "async"},
)
print(r.status_code, r.json())

What should I verify before relying on it?

The model-by-model audio list names which ids take audio on Sume.

  • Read supported_input_references for the model id at GET /v1/videos/models; only listed types are accepted.
  • Confirm the audio is WAV or MP3, 15 MB or smaller and 15 seconds total.
  • Do not use a reference audio when you want Sume to generate speech from the prompt; that is a different input.
  • Remember MiniMax's own rule that H3 reference audio is a reference, not a lip-sync track.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume