Wan 2.7 custom audio (2-30 s) vs Sume Wan 3.0 audio references
Alibaba's Wan 2.7 takes a custom WAV or MP3 of 2-30 seconds. Sume's Wan 3.0 accepts up to 5 reference audio clips, 15 seconds in total, with an image or video.

If you used Wan 2.7's custom audio input at Alibaba, the nearest Sume setting is reference_audio_urls on wan-3.0, but the limits differ: Sume takes at most 5 audio clips totalling 15 seconds, and always needs an image or video reference next to them. Alibaba's Wan 2.7 text-to-video page lists a custom audio file of WAV or MP3, 2 to 30 seconds, up to 15 MB (Alibaba Model Studio, read 2026-10-10).
So a 30-second narration track from a Wan 2.7 job cannot be sent whole to Sume's Wan 3.0 row.
What Alibaba lists for Wan 2.7
The page describes 2 to 15 second clips at 720P or 1080P, aspect ratios 16:9, 9:16, 1:1, 4:3 and 3:4, MP4 with H.264, a typical wait of 1 to 5 minutes, negative prompts, a watermark option and automatic background music or effects. It also warns that the task id and video URL are valid for only 24 hours.
What Sume lists for Wan 3.0
Sume's Wan 3.0 row accepts 2 to 30 seconds at 480p, 720p and 1080p, with 16:9, 4:3, 1:1, 3:4, 9:16 and the auto and adaptive ratios. The catalog caps references at 10 images, 5 videos and 5 audio clips; reference videos must be at most 15 seconds in total and at least 16 fps, and reference audio at most 15 seconds in total. It has no bitrate_mode, and file_url, web_url and enable_thinking are not exposed in v1.
| Item | Alibaba Wan 2.7 | Sume Wan 3.0 |
|---|---|---|
| Audio input | Custom WAV or MP3 | reference_audio_urls |
| Length per file | 2 to 30 s | Not stated per file |
| Total audio length | One file | At most 15 s in total |
| Number of files | One | At most 5 |
| Needs other references | Not stated | Yes: an image or a video |
| Clip length | 2 to 15 s | 2 to 30 s |
Port plan for a long narration track
If your soundtrack is longer than 15 seconds, split it to fit and keep the clips short.
- Cut the track into segments of 15 seconds or less and submit one job per segment.
- Send at least one reference image next to the audio; audio alone is rejected.
- Join the finished clips on a timeline and lay the full track under them if you need one continuous voice.
- Fetch each result with your bearer key from the content endpoint; Sume's jobs docs describe
unsigned_urlsas the download source.
A note on what audio does
A reference audio clip on Sume is guidance for the model, not a guarantee of lip-sync. For tight speech-to-mouth timing, use a dedicated lip-sync or avatar route, and treat Wan's output as picture with sound.
Sources
Related posts
More in Comparisons
- Canva Connect API vs the Sume API: what each is for
Canva Connect syncs designs, assets and comments, with some APIs in preview. The Sume API generates video, images and audio. Different jobs.
- CapCut lists Seedance 2.5 and Gemini Omni. Does Sume's API carry them?
CapCut's tools page names several video and image models. Sume's API documents Seedance 2.5 and Gemini Omni Flash 1.1; the rest are not in its docs.
- ChatGPT Image 2 vs 2.5 on Sume: $0.26375 vs $0.065875 and what differs
On Sume, GPT Image 2.5 high quality at 1024 costs $0.065875, a quarter of GPT Image 2 at $0.26375, and adds mask_url, background and 16 references.
- Creatomate RenderScript vs a Sume Timeline document: field map
Creatomate's RenderScript is a general scene JSON; Sume's Timeline 1.0 is one audio spine plus video slots. Field-by-field map and what Sume cannot express.
Written by Sume