Recipe blog post to video: scrape text, send step photos to Wan 3.0
Wan 3.0 can take a webpage as a reference at Alibaba, but Sume does not expose web_url. Scrape the recipe, then send step photos as references. Steps and curl.

To turn a recipe blog post into a video with Wan 3.0 on Sume, do not paste the URL into the video request. Read the page first, with the hosted MCP tool crawl_scrape, write a shot-by-shot prompt from the ingredients and steps, and send three to six of the post's step photos as reference_image_urls to wan-3.0. Alibaba's Wan 3.0 README (read 2026-10-05) lists webpages among the up to 20 reference assets the model accepts. Sume's catalog entry for wan-3.0 states that file_url / web_url are not exposed in v1, so on Sume the page has to be turned into text and images before it reaches the model.
What to take from the page
A recipe post has four things worth carrying over: a title, an ingredient list, numbered steps, and photos. The video should show the steps, not recite the ingredient list, so keep the ingredients for a caption or a pinned comment.
Scraping is the step that replaces Wan 3.0's native webpage input. crawl_scrape is one of the read tools in the hosted MCP inventory; an agent in Claude or Cursor can call it, summarize the page into the prompt, and hand the photo URLs to the video request. Do the same by hand if you prefer: copy the steps into a text file and the photo URLs into a list. Either way, the model never sees the live page, so what is not in your prompt or references will not be in the clip.
Keep the whole prompt under what the model can follow in one pass. Five steps in twenty seconds is a comfortable rhythm; fourteen steps in twenty seconds is not, and the clip will skip some.
- Title and yield: goes into the first line of the prompt so the dish is named once.
- Steps: shorten each to one visible action ("fold the egg whites in"), one action per shot.
- Photos: pick the photo that shows the state after the step. These become the references.
- Ingredient quantities: keep them out of the prompt. Generated on-screen text can drift, so add exact amounts later with captions.
How the post maps to a Wan 3.0 request
Sume's Video generation docs say Wan 3.0 accepts image, video and audio references, and wan-3.0 accepts 2 to 30 seconds. The catalog constraint for the model lists at most 10 reference images. A recipe with six step photos fits in one request.
| Page element | Request field | Note |
|---|---|---|
| Dish name and mood | prompt | One sentence, then the steps in order |
| Step photos (3 to 6) | reference_image_urls | Public https URLs; at most 10 images for wan-3.0 |
| Total length | duration | 2 to 30 seconds; about 4 seconds per step is a start |
| Feed format | aspect_ratio | 9:16 for Reels and Shorts, 16:9 for the blog embed |
| Retry safety | Idempotency-Key header | One key per attempt you intend to pay for |
Example request
This request asks for a 20 second vertical clip from five step photos. The endpoint, model, duration, aspect_ratio and mode fields are the ones in the Video Router docs; the reference field names follow the reference_*_urls pattern the same page uses.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: recipe-001" \
-d '{
"model": "wan-3.0",
"prompt": "A home cook makes lemon pancakes in five steps: whisk, fold, pour, flip, stack. Warm kitchen light, hands only, no text on screen.",
"reference_image_urls": ["https://example.com/step1.jpg", "https://example.com/step2.jpg", "https://example.com/step3.jpg", "https://example.com/step4.jpg", "https://example.com/step5.jpg"],
"resolution": "720p",
"duration": 20,
"aspect_ratio": "9:16",
"mode": "async"
}'Check before you publish
Job status is async: poll the returned job, or set a callback, then fetch the clip. Before posting, sample a few stills with video inspect to confirm the food in each shot matches the step. If the steps look merged into one motion, split the job: one request per two steps, joined with Timeline 1.0.
If the recipe needs exact quantities on screen, burn them as captions with video captions cues rather than asking the model to draw the numbers.
Sources
Related posts
More in Use cases
- Recolor a product with a swatch reference and mask on GPT Image 2.5
BFL's FLUX 3 Image lists recolor, replace and move edits by bounding box. On Sume, a recolor job takes a swatch reference and a mask_url on GPT Image 2.5.
- Redact names from a transcript and the audio: STT words, then split
Sume STT returns word timings with a type field. Find the names you list, blank them in text, and cut the same spans from the audio for $0.01.
- Reels ad disclaimer: keep the bottom 40% clear at 1080x1920
Meta says to leave the bottom 40% of a Reels ad free of text and logos when it carries a disclaimer. Pixel math and where to burn captions with Sume.
- Reels ads ban GIFs, face effects and licensed music: AI clip preflight
Meta's Reels ad spec says no licensed music, face or camera effects, GIFs or product tags. A five-point preflight for an AI-made 9:16 clip before upload.
Written by Sume