A 50-page deck to AI video: five ten-slide Wan 3.0 batches on Sume

Alibaba caps document input at 50 pages. Sume wan-3.0 takes 10 reference images per job, so a 50-slide deck becomes five jobs. Cost at 6 s or 30 s per batch.

5 min readSume
All posts

A 50-page deck becomes five wan-3.0 jobs of ten slide images each on Sume. Alibaba's page lets you send one document of up to 50 pages and 100 MB to its own Wan 3.0 endpoint. Sume's request validation limits wan-3.0 to 10 reference images per job, and its docs list no document type, so you export slides to PNG and batch them yourself.

Why ten is the number

The cap is set in the API's schema validation: wan-3.0 accepts up to 10 images, 5 videos and 5 audio references per request. Fifty slides divided by ten is five jobs. Each job needs a prompt that says what the viewer should see for those ten slides, in order.

Wan 3.0 cost for a 50-slide deck in five batches on Sume (read 2026-10-07)
Plan480p720p1080p
Five 6 s batches (30 s total)$1.90$3.75$7.50
Five 30 s batches (150 s total)$9.40$18.75$37.50

Arithmetic

At 720p a 6 s batch is 6 x $0.10 x 1.25 = $0.75, so five batches are $3.75, the same as one 30 s job. A 30 s batch is $3.75, so five are $18.75. At 1080p the five 30 s batches are $37.50. Pick the duration from the content: ten slides in 6 s is under a second per slide, while 30 s gives three seconds each.

Caveats that change the output

  • Sume describes reference images as visual guidance, not exact frames. Slide text may be redrawn, and Alibaba lists on-screen text accuracy as still improving.
  • Use first_frame and last_frame if two slides must appear exactly, but if frame_images and references are both sent, frame_images controls the mode.
  • Batches are independent jobs, so keep the same prompt style and aspect ratio across all five to avoid visible cuts at the joins.
  • Pass an Idempotency-Key for each batch so a retry does not bill a second job.

Building the five requests

This script splits fifty slide URLs into five payloads. Post each one to /v1/videos with its own Idempotency-Key.

import json

slides = [f"https://example.com/deck/slide-{n:02d}.png" for n in range(1, 51)]
batches = [slides[i:i + 10] for i in range(0, len(slides), 10)]
jobs = []
for n, batch in enumerate(batches, start=1):
    jobs.append({
        "model": "wan-3.0",
        "prompt": f"Part {n} of 5: walk through these slides in order.",
        "duration": 6,
        "resolution": "720p",
        "input_references": [
            {"type": "image_url", "image_url": {"url": u}} for u in batch
        ],
    })
print(len(jobs), "jobs,", len(jobs[0]["input_references"]), "images each")
print(json.dumps(jobs[0], indent=2)[:200])

Keeping the five parts consistent

Five independent jobs will not share a look unless you make them. Reuse the same style sentence in every prompt, such as a flat motion-graphics look with a fixed colour palette, and keep resolution and aspect_ratio identical. Put a slide that works as an opening frame at the start of each batch so that every part starts from a similar composition.

Order matters inside a batch. The prompt should name the sequence in plain words, for example opening with the title slide, then the market slide, then the three product slides. The model is not told slide numbers by any documented syntax, so describe the content rather than writing slide 4.

What only the vendor page promises

Alibaba says its document input treats your prompt as the control centre and can turn a training deck into video courseware or a product deck into a launch ad. Those outcomes are described for Alibaba's endpoint. On Sume you get the reference-image route above, which gives you control over which slides go in which job, but not a document parser.

Shorter option

If you only need highlights, pick ten key slides and make one job. A single 15 s clip at 720p costs 15 x $0.125 = $1.875, which Sume rounds to $1.88.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume