Wan 3.0 auto scene splitting vs explicit scene slots in Sume Timeline
Wan 3.0 advertises auto scene splitting inside one generation. If you want control of each scene, build separate clips and join them with Timeline 1.0.

The Wan 3.0 README (read 2026-10-05) lists auto scene splitting among its features, next to native 30-second generation. In plain terms the model can cut a long prompt into scenes on its own. If you want to decide where the cuts fall, how long each scene runs and which transition joins them, generate the scenes as separate clips and assemble them with Timeline 1.0. The trade is control against a single request.
Two ways to a 30-second multi-scene video
On Sume, wan-3.0 accepts 2 to 30 seconds, so both routes are open to you. Pick the single request when the scenes are loose and cheap to redo; pick the slots when a client will review scene by scene.
| Question | One Wan 3.0 request | Several clips plus Timeline 1.0 |
|---|---|---|
| Who picks the cut points | The model | You |
| Retry one bad scene | Re-run the whole generation | Re-run only that clip |
| Transitions | Whatever the model renders | fade, wipeleft, wiperight, slideup, slidedown, dissolve |
| Mix models per scene | No | Yes, each slot can come from a different clip |
| Number of generation requests | 1 | One per scene |
Build it with slots
Timeline 1.0 takes one audio spine plus ordered video[] slots and returns one MP4. The required fields are audio.duration_seconds and 1 to 200 video slots, every URL on your workspace's media.sume.com. The first slot must start at 0, later starts must increase, and a transition may only appear on slots after the first. If you have no voice track yet, audio.mode: "silence" declares the length without a spine file.
Run POST /v1/timeline-1.0/plan first. It is unbilled, creates no job, and returns duration_seconds, segment_count and estimated_cost_usd_micros, which tells you the compile is valid before you spend on renders.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scenes-001" \
-d '{
"audio": { "mode": "silence", "duration_seconds": 24 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/scene1.mp4", "start": 0, "duration": 8 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/scene2.mp4", "start": 8, "duration": 8,
"transition": { "type": "fade", "duration": 0.25 } },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/scene3.mp4", "start": 16, "duration": 8,
"transition": { "type": "dissolve", "duration": 0.25 } }
]
}'When the single request is the better call
Explicit slots are not always worth it. For a mood piece, a background loop or a concept test, one request is simpler: one prompt, one job id, one MP4, and nothing to assemble. The cost of a bad outcome is a re-run of one job. Choose explicit slots when a reviewer will sign off scene by scene, when different scenes need different references, or when one scene must be redone without touching the rest.
A good compromise is to generate once, review, and only split the video into slots if a scene fails. Video trim cuts a [start, end) range from the finished clip for $0.02 per job, so a failed scene can be replaced by a regenerated clip placed into the same slot of the timeline.
- Mood piece or loop: one request.
- Client-reviewed storyboard: separate clips.
- Mixed models per scene: separate clips.
Naming and storage
Name each scene clip by its order, such as scene-01 and scene-02, and keep the slot start and duration beside the name. When the timeline compiles, the declared starts are authoritative, so a spreadsheet of starts and durations is the plan. Import every clip to media.sume.com before the render, because the render accepts only your workspace's hosted artifacts.
Cost of the join
The Timeline docs put the public rate at $0.10 per output minute, rounded up, with no provider inference, only worker ffmpeg. The generation requests are the main cost; the join is small. Confirm the live rate in GET /v1/catalog.
Sources
Related posts
More in Use cases
- Wan 3.0: one 30-second prompt or three 10-second clips on Sume
Alibaba says Wan 3.0 splits scenes itself. On Sume, one 30 s request and three 10 s clips cost $3.75 at 720p each way; with three clips a redo costs $1.25.
- Twelve sequential images then animate: Wan 3.0 storyboard cost
Alibaba says Wan 3.0 makes up to 12 sequential images. On Sume, make stills with an image model, then animate each: 12 two-second 480p clips cost $1.50.
- Wan 3.0 proof clip: test on-screen text before a 30-second render
A 2 second 480p Wan 3.0 clip costs $0.125 on Sume. Use it to test on-screen text per language before a 30 second render that costs $1.875 to $7.50.
- Wave.video Streamer $16 vs weekly restaurant promos on Sume
Wave.video lists Streamer $16, Creator $24 and Business $48. A 15-second weekly special from a phone clip and a voice line costs about $0.11 on Sume.
Written by Sume