90-second YouTube Short from three Wan 3.0 clips and a voice file
Build a 90-second vertical Short from three 30-second Wan 3.0 clips, one voice track and one Timeline render. Request shapes and the full cost at 480p to 1080p.

A 90-second Short is three 30-second Wan 3.0 clips laid on one voice track by a single Timeline 1.0 render. At 720p that is $11.45 before voice, and the render itself costs $0.20. YouTube's help page says Shorts can run up to three minutes in square or vertical (read 2026-10-08), so 90 seconds sits well inside the limit.
Wan 3.0 matters here because one request reaches 30 seconds. Alibaba lists wan3.0-video at 2 to 30 seconds with 480P, 720P and 1080P output (read 2026-10-08), and Sume's docs give wan-3.0 the same 2 to 30 second range. Three requests cover the whole Short with no stitching inside a shot.
Step 1: three clip requests
Each clip is an ordinary asynchronous video job. Use one Idempotency-Key per clip so a retry returns the original job.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: short90-clip-1" \
-d '{
"model": "wan-3.0",
"prompt": "Slow push-in on a ceramics studio, hands shaping a bowl",
"duration": 30,
"resolution": "720p",
"aspect_ratio": "9:16"
}'Step 2: one render
The Timeline body has one audio spine and ordered video[] slots. audio.duration_seconds is the output length, and slot starts must increase from 0. Every URL must already be a media.sume.com artifact of your workspace, which generated clips are.
{
"audio": { "url": "https://media.sume.com/artifacts/artf_demo/voice.wav", "duration_seconds": 90 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/c1.mp4", "start": 0, "duration": 30 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/c2.mp4", "start": 30, "duration": 30,
"transition": { "type": "fade", "duration": 0.25 } },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/c3.mp4", "start": 60, "duration": 30,
"transition": { "type": "fade", "duration": 0.25 } }
]
}What it costs
Voice at 1,500 characters is added at Sume's TTS rate of $0.0475 per 1,000 characters. Resolution is the only variable.
| Resolution | 3 clips | Voice (1,500 chars) | Timeline render | Total |
|---|---|---|---|---|
| 480p | $5.625 | $0.0713 | $0.20 | $5.8963 |
| 720p | $11.25 | $0.0713 | $0.20 | $11.5212 |
| 1080p | $22.50 | $0.0713 | $0.20 | $22.7712 |
Practical notes
- The default output is 1080 by 1920, which is the vertical case. Square needs
output.widthandoutput.height. - Keep the voice length equal to the sum of slot durations. Coverage may stop at most 0.5 seconds before the spine ends.
- Draft at 480p, then rerun only the approved prompts at 1080p.
Sequencing and consistency
Generate clip one first and read it before you spend on the others. Wan 3.0 clips are independent jobs, so a character or place in clip one will not carry into clip two unless the prompts describe it the same way. Reuse one sentence for the look (lens, light, palette) across all three prompts, and vary only the action. If a clip fails or is rejected, resubmit just that slot; the other two files are untouched, and the Timeline body only needs its URL replaced. Poll each job with GET /v1/jobs/:id/status, or register a webhook so the render starts the moment the third clip is ready.
Sources
Related posts
More in Use cases
- Ad hook test matrix: 4 hooks x 2 ratios on Omni Flash, priced
Four ad hooks in 16:9 and 9:16 as 5-second Omni Flash drafts cost $1.50 at 360p on Sume; re-rendering two winners at 1080p adds $1.875. The full matrix.
- Roleplay team budget: Synthesia learner seats or a Sume clip library
Twelve-month cost of Synthesia Roleplay Training for 20 learners against a library of 24 scenario clips on Sume, and where each stops being the right tool.
- AI children's story video: 8 Omni scenes for $8.00
Eight 8-second Omni scenes at $1.00 each make a one-minute story on Sume. Use drawn characters: Google blocks minors in image edits in the EEA, UK and CH.
- AI dubbing workflow on Sume: detach, transcribe, TTS, timeline
Sume has no one-call dubbing endpoint. Chain detach, STT, your translation, TTS and a timeline render for a 60-second video at about $0.16 plus translation.
Written by Sume