Restaurant menu PDF to a vertical video: Wan 3.0 with pages as images
A menu PDF cannot go into a Sume video request. Export the pages as images, send up to 10 as Wan 3.0 references, and add the prices as captions. Steps.

To make a short video from a restaurant menu PDF with Wan 3.0 on Sume, export the PDF pages to images, pick the dishes you want to show, and send those images as references to wan-3.0. Alibaba's Wan 3.0 README (read 2026-10-05) says the model accepts up to 20 reference assets, documents included. Sume's wan-3.0 catalog entry does not expose file_url or web_url in v1, and its reference field for images is reference_image_urls with at most 10 entries. So the PDF is a source you prepare, not a file you upload.
Why not send the whole menu
A menu page is mostly small text. A video model that is shown a page of dish names and prices will treat it as texture, and the output can show invented or garbled item names. The README lists text rendering in 12 languages as a feature, but a menu is the case where a wrong digit costs you: a $14 dish shown as $41 is a real error. Treat any price the model draws as unverified.
The same applies to dish names in a script the model cannot read at small size, such as a handwritten specials board. If a name matters, it should come from your text layer, not from pixels the model re-draws.
The safer split is to let the model make the food footage and let your own text carry the facts.
A workable pipeline
Three steps cover it. Export each menu page, or better each dish photo, as a JPEG or PNG and host it at an https URL. Pick at most ten images, one per dish you want on screen. Submit one wan-3.0 request per section of the menu (starters, mains, desserts) so each clip stays short and focused.
| Menu element | Where it goes | Reason |
|---|---|---|
| Dish photos | reference_image_urls on wan-3.0 | Up to 10 images per request |
| Section mood (warm, fast, steam) | prompt | One sentence per dish, in order |
| Prices and dish names | video captions cues | You type the text; the model does not |
| Clip length | duration (2 to 30) | Roughly 3 to 4 seconds per dish |
| Frame shape | aspect_ratio 9:16 | Vertical for Reels, Shorts and TikTok |
Burn the prices in yourself
After the clip finishes, use video captions. The page says you can send cues (or segments) with text, start and end to burn authored overlay text without speech-to-text, so a price list appears at the seconds you choose. That keeps the numbers exactly as on the PDF.
The captions input is a public HTTPS video URL, so use the finished clip's content URL or your own host.
Run the caption job after you have chosen the final cut of the clip. A caption burn is a separate job, so a re-render of the video means burning again. Plan one caption pass per approved clip, not per draft.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"video_url": "https://example.com/mains-clip.mp4",
"cues": [
{"text": "Brisket plate $18", "start": 0.5, "end": 3.5},
{"text": "Smoked wings $12", "start": 3.5, "end": 7.0}
]
}'Limits to plan around
The reference cap of 10 images is per request, not per menu. A twenty-dish menu is two or three clips joined with Timeline 1.0, which takes an audio spine and ordered video slots and returns one MP4. Allergen and dietary labels should never come from generated text: type them as cues and check them by eye before posting.
One more habit helps: keep the source PDF page number next to each image URL in your notes. When a viewer asks why a dish looks different from the menu, you can trace the clip back to the exact photo and prompt that produced it, and regenerate only that section.
Sources
Related posts
More in Use cases
- Room restyle with a moodboard on Ideogram 4.5: photo goes first
On Sume the first input_references image is the room being edited and later ones are style references. Send photo, then moodboard, and omit aspect_ratio.
- SaaS app UGC video on Seedance 2.5: screen reference and cost
Make a UGC-style SaaS app video on Seedance 2.5 from a product screenshot: what a screen reference can and cannot do, prompt, and cost for 10 to 20 seconds.
- SaaS release teaser: a 15-second Wan 3.0 clip at 720p vs 1080p
A 15-second teaser from a UI screenshot with wan-3.0 on Sume costs $1.875 at 720p and $3.75 at 1080p; a draft at 480p is $0.9375.
- Safety module in 10 languages: 60 seconds, 10 caption jobs, $2.00
A 60-second safety module in ten languages is ten Sume caption jobs at $0.20, $2.00, plus about two cents for one transcript. Index-Translate does the text.
Written by Sume