Restaurant menu PDF to a vertical video: Wan 3.0 with pages as images

A menu PDF cannot go into a Sume video request. Export the pages as images, send up to 10 as Wan 3.0 references, and add the prices as captions. Steps.

5 min readSume
All posts

To make a short video from a restaurant menu PDF with Wan 3.0 on Sume, export the PDF pages to images, pick the dishes you want to show, and send those images as references to wan-3.0. Alibaba's Wan 3.0 README (read 2026-10-05) says the model accepts up to 20 reference assets, documents included. Sume's wan-3.0 catalog entry does not expose file_url or web_url in v1, and its reference field for images is reference_image_urls with at most 10 entries. So the PDF is a source you prepare, not a file you upload.

Why not send the whole menu

A menu page is mostly small text. A video model that is shown a page of dish names and prices will treat it as texture, and the output can show invented or garbled item names. The README lists text rendering in 12 languages as a feature, but a menu is the case where a wrong digit costs you: a $14 dish shown as $41 is a real error. Treat any price the model draws as unverified.

The same applies to dish names in a script the model cannot read at small size, such as a handwritten specials board. If a name matters, it should come from your text layer, not from pixels the model re-draws.

The safer split is to let the model make the food footage and let your own text carry the facts.

A workable pipeline

Three steps cover it. Export each menu page, or better each dish photo, as a JPEG or PNG and host it at an https URL. Pick at most ten images, one per dish you want on screen. Submit one wan-3.0 request per section of the menu (starters, mains, desserts) so each clip stays short and focused.

Menu to Sume mapping (read 2026-10-05)
Menu elementWhere it goesReason
Dish photosreference_image_urls on wan-3.0Up to 10 images per request
Section mood (warm, fast, steam)promptOne sentence per dish, in order
Prices and dish namesvideo captions cuesYou type the text; the model does not
Clip lengthduration (2 to 30)Roughly 3 to 4 seconds per dish
Frame shapeaspect_ratio 9:16Vertical for Reels, Shorts and TikTok

Burn the prices in yourself

After the clip finishes, use video captions. The page says you can send cues (or segments) with text, start and end to burn authored overlay text without speech-to-text, so a price list appears at the seconds you choose. That keeps the numbers exactly as on the PDF.

The captions input is a public HTTPS video URL, so use the finished clip's content URL or your own host.

Run the caption job after you have chosen the final cut of the clip. A caption burn is a separate job, so a re-render of the video means burning again. Plan one caption pass per approved clip, not per draft.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "video_url": "https://example.com/mains-clip.mp4",
    "cues": [
      {"text": "Brisket plate  $18", "start": 0.5, "end": 3.5},
      {"text": "Smoked wings  $12", "start": 3.5, "end": 7.0}
    ]
  }'

Limits to plan around

The reference cap of 10 images is per request, not per menu. A twenty-dish menu is two or three clips joined with Timeline 1.0, which takes an audio spine and ordered video slots and returns one MP4. Allergen and dietary labels should never come from generated text: type them as cues and check them by eye before posting.

One more habit helps: keep the source PDF page number next to each image URL in your notes. When a viewer asks why a dish looks different from the menu, you can trace the clip back to the exact photo and prompt that produced it, and regenerate only that section.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume