Recipe blog post to video: scrape text, send step photos to Wan 3.0

Wan 3.0 can take a webpage as a reference at Alibaba, but Sume does not expose web_url. Scrape the recipe, then send step photos as references. Steps and curl.

5 min readSume
All posts

To turn a recipe blog post into a video with Wan 3.0 on Sume, do not paste the URL into the video request. Read the page first, with the hosted MCP tool crawl_scrape, write a shot-by-shot prompt from the ingredients and steps, and send three to six of the post's step photos as reference_image_urls to wan-3.0. Alibaba's Wan 3.0 README (read 2026-10-05) lists webpages among the up to 20 reference assets the model accepts. Sume's catalog entry for wan-3.0 states that file_url / web_url are not exposed in v1, so on Sume the page has to be turned into text and images before it reaches the model.

What to take from the page

A recipe post has four things worth carrying over: a title, an ingredient list, numbered steps, and photos. The video should show the steps, not recite the ingredient list, so keep the ingredients for a caption or a pinned comment.

Scraping is the step that replaces Wan 3.0's native webpage input. crawl_scrape is one of the read tools in the hosted MCP inventory; an agent in Claude or Cursor can call it, summarize the page into the prompt, and hand the photo URLs to the video request. Do the same by hand if you prefer: copy the steps into a text file and the photo URLs into a list. Either way, the model never sees the live page, so what is not in your prompt or references will not be in the clip.

Keep the whole prompt under what the model can follow in one pass. Five steps in twenty seconds is a comfortable rhythm; fourteen steps in twenty seconds is not, and the clip will skip some.

  • Title and yield: goes into the first line of the prompt so the dish is named once.
  • Steps: shorten each to one visible action ("fold the egg whites in"), one action per shot.
  • Photos: pick the photo that shows the state after the step. These become the references.
  • Ingredient quantities: keep them out of the prompt. Generated on-screen text can drift, so add exact amounts later with captions.

How the post maps to a Wan 3.0 request

Sume's Video generation docs say Wan 3.0 accepts image, video and audio references, and wan-3.0 accepts 2 to 30 seconds. The catalog constraint for the model lists at most 10 reference images. A recipe with six step photos fits in one request.

Recipe part to Sume request field (read 2026-10-05)
Page elementRequest fieldNote
Dish name and moodpromptOne sentence, then the steps in order
Step photos (3 to 6)reference_image_urlsPublic https URLs; at most 10 images for wan-3.0
Total lengthduration2 to 30 seconds; about 4 seconds per step is a start
Feed formataspect_ratio9:16 for Reels and Shorts, 16:9 for the blog embed
Retry safetyIdempotency-Key headerOne key per attempt you intend to pay for

Example request

This request asks for a 20 second vertical clip from five step photos. The endpoint, model, duration, aspect_ratio and mode fields are the ones in the Video Router docs; the reference field names follow the reference_*_urls pattern the same page uses.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: recipe-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "A home cook makes lemon pancakes in five steps: whisk, fold, pour, flip, stack. Warm kitchen light, hands only, no text on screen.",
    "reference_image_urls": ["https://example.com/step1.jpg", "https://example.com/step2.jpg", "https://example.com/step3.jpg", "https://example.com/step4.jpg", "https://example.com/step5.jpg"],
    "resolution": "720p",
    "duration": 20,
    "aspect_ratio": "9:16",
    "mode": "async"
  }'

Check before you publish

Job status is async: poll the returned job, or set a callback, then fetch the clip. Before posting, sample a few stills with video inspect to confirm the food in each shot matches the step. If the steps look merged into one motion, split the job: one request per two steps, joined with Timeline 1.0.

If the recipe needs exact quantities on screen, burn them as captions with video captions cues rather than asking the model to draw the numbers.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume