Wan 2.7 custom audio (2-30 s) vs Sume Wan 3.0 audio references

Alibaba's Wan 2.7 takes a custom WAV or MP3 of 2-30 seconds. Sume's Wan 3.0 accepts up to 5 reference audio clips, 15 seconds in total, with an image or video.

4 min readSume
All posts

If you used Wan 2.7's custom audio input at Alibaba, the nearest Sume setting is reference_audio_urls on wan-3.0, but the limits differ: Sume takes at most 5 audio clips totalling 15 seconds, and always needs an image or video reference next to them. Alibaba's Wan 2.7 text-to-video page lists a custom audio file of WAV or MP3, 2 to 30 seconds, up to 15 MB (Alibaba Model Studio, read 2026-10-10).

So a 30-second narration track from a Wan 2.7 job cannot be sent whole to Sume's Wan 3.0 row.

What Alibaba lists for Wan 2.7

The page describes 2 to 15 second clips at 720P or 1080P, aspect ratios 16:9, 9:16, 1:1, 4:3 and 3:4, MP4 with H.264, a typical wait of 1 to 5 minutes, negative prompts, a watermark option and automatic background music or effects. It also warns that the task id and video URL are valid for only 24 hours.

What Sume lists for Wan 3.0

Sume's Wan 3.0 row accepts 2 to 30 seconds at 480p, 720p and 1080p, with 16:9, 4:3, 1:1, 3:4, 9:16 and the auto and adaptive ratios. The catalog caps references at 10 images, 5 videos and 5 audio clips; reference videos must be at most 15 seconds in total and at least 16 fps, and reference audio at most 15 seconds in total. It has no bitrate_mode, and file_url, web_url and enable_thinking are not exposed in v1.

Audio input: Alibaba page (read 2026-10-10) and Sume Wan 3.0 catalog in the repo on 2026-10-10
ItemAlibaba Wan 2.7Sume Wan 3.0
Audio inputCustom WAV or MP3reference_audio_urls
Length per file2 to 30 sNot stated per file
Total audio lengthOne fileAt most 15 s in total
Number of filesOneAt most 5
Needs other referencesNot statedYes: an image or a video
Clip length2 to 15 s2 to 30 s

Port plan for a long narration track

If your soundtrack is longer than 15 seconds, split it to fit and keep the clips short.

  • Cut the track into segments of 15 seconds or less and submit one job per segment.
  • Send at least one reference image next to the audio; audio alone is rejected.
  • Join the finished clips on a timeline and lay the full track under them if you need one continuous voice.
  • Fetch each result with your bearer key from the content endpoint; Sume's jobs docs describe unsigned_urls as the download source.

A note on what audio does

A reference audio clip on Sume is guidance for the model, not a guarantee of lip-sync. For tight speech-to-mouth timing, use a dedicated lip-sync or avatar route, and treat Wan's output as picture with sound.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume