Three-minute YouTube Short from six 30 s AI clips: Wan vs Seedance
YouTube counts square or vertical videos up to three minutes as Shorts. Six 30 s 9:16 clips cost $22.50 on Wan 3.0 at 720p or $104.04 on Seedance 2.5.

Six 30-second vertical clips make a three-minute Short, and at 720p they cost $22.50 on Wan 3.0 or $104.04 on Seedance 2.5 on Sume. YouTube's help page says videos uploaded after October 15, 2024 with a square or vertical aspect ratio up to three minutes long are categorized as Shorts on standard channels, so the length cap lines up with six jobs of the 30 s model maximum.
| Plan (6 x 30 s, 9:16) | 480p | 720p | 1080p |
|---|---|---|---|
| Wan 3.0 | $11.28 | $22.50 | $45.00 |
| Seedance 2.5 | $48.42 | $104.04 | $255.90 |
Arithmetic
Wan 3.0 at 720p is 30 x $0.10 x 1.25 = $3.75 a job, so six jobs are $22.50. Seedance 2.5 at 720p in 9:16 is 720 x 1280, so 648,000 tokens and $17.34 a job, and six jobs are $104.04. Both models accept 9:16 and 30 s. At 1080p Wan is $45.00 and Seedance is $255.90 (6 x $42.65).
What YouTube says that affects the plan
- Square or vertical only. A 16:9 render would not count as a Short.
- A Short longer than one minute with an active Content ID claim is blocked globally, according to the help page, so generated audio should not borrow a copyrighted track.
- The help page says most Shorts Audio Library songs can be used for up to 90 seconds in a three-minute Short, so a 180 s film needs its own sound bed for the rest.
- The page does not mention AI-generated video, so rules for that are outside this post.
Why six and not one
Neither model produces a three-minute clip in one pass. Sume's docs give 30 s as the top of the range for wan-3.0 and seedance-2.5, and the 2 to 30 s and 4 to 30 s ranges show where each starts. So the three-minute film is a sequence by construction, and the question is only how to hide the six joins. Hold a stable shot at each end, avoid mid-motion cuts, and let each part open on the frame the previous one closed with.
Another approach is to shape the film so that the joins look like scene changes: six distinct locations, one per part, with a cut at each boundary. That needs no frame handoff and lets you generate the six parts in parallel, since no part depends on another.
Audio across six parts
Each job generates or omits its own audio, so six parts mean six separate sound beds with audible seams at 30 second marks. Either set generate_audio to false throughout and lay one music track under the joined film in a timeline, or accept the seams as scene changes. Mind the YouTube Content ID note above when choosing the track.
Keeping six parts together
Each job is independent. Reuse the same style sentence in every prompt and hand off a last frame from each part as the first frame of the next, using video-frames then frame_images. A 30 s clip gives you a clean cut point every 30 seconds, so write the script in six beats and let each end on a held shot.
Draft at 480p first: six Wan 3.0 drafts are 6 x $1.88 = $11.28, which tells you whether the script works before $22.50 or more is spent.
Sources
Related posts
More in Use cases
- How do I put three products in one AI video from reference photos?
Send three product photos as input_references to Gemini Omni Flash 1.1: one 8-second 720p clip is $1.00 on Sume, against $1.89 for three single clips.
- TikTok TopView safe zone: check the open-screen and in-feed frames
TikTok's TopView page has separate safe zones for the open screen and the in-feed stage. Pull a frame from each moment with Sume video inspect and check it.
- How do I make a trade-show booth audio loop with AI voice and music?
Make a 60-second booth loop from one TTS job, one music bed and a Timeline render with a looping soundtrack: about 25 cents on Sume, then play it on repeat.
- Transcribe an interview with timestamps via API (no speaker labels)
Send interview audio to Sume STT for text, word start and end times and sentence segments. It returns no speaker names. 25 minutes costs 28 cents.
Written by Sume