Wan 3.0 references or start/end frames, not both: two Sume calls
Krea says a Wan 3.0 job takes reference media or start/end frames, not both. Sume lists them as separate modes, so split the work into two requests.

Krea's Wan 3.0 entry says a single job accepts reference media or start and end frames, not both. Sume's video docs also list those as separate input modes, so plan two requests: one guided by a start frame, one guided by references. On POST /v1/videos the docs say that if both frame_images and input_references are sent, frame_images takes precedence and the request is treated as image-to-video, so the references are not used; the video input-modes page does not describe a mix, so do not rely on one.
Krea facts are from its changelog (August 28, 2026); Sume facts are from the video input modes and video generation docs, read 2026-10-01.
What is the vendor rule?
Krea says Wan 3.0 can generate from a text prompt, animate a start image, or be guided by reference images, videos, and audio, at 480p, 720p, or 1080p. For a job, you attach reference media or you set start and end frames.
How does Sume list the modes?
Each mode is its own row in the docs, with its own fields.
| Mode | Fields |
|---|---|
| Image to video | image_url (first frame) |
| Start + end frames | image_url + end_image_url |
| Reference-guided | reference_image_urls and/or reference_video_urls, optional reference_audio_urls |
What can I not send?
end_image_url requires image_url, so an end frame alone is invalid. reference_audio_urls requires at least one reference image or video. Wan 3.0 is among the models the docs list as honoring audio and video references; its duration range is 2 to 30 seconds.
How do I split one idea into two requests?
Pick the control that matters most for each clip. Use start and end frames when you need a specific opening and closing image. Use references when you need a character or style held across a shot. For a clip that needs both, make the reference-guided clip first and then use a frame from it as image_url on the second request; see frame images and input references together for what the docs say about combining them.
Sources
Related posts
More in Models
- LTX-2.5 Cinemagraph LoRA vs image-to-video on Sume
The LTX-2.5 Cinemagraph LoRA is an image-to-video adapter for selective motion. On Sume, start from first_frame and describe the motion in the prompt.
- Luma layers API: 10 RGBA layers vs Sume's flat image result
Luma type layering splits one image into up to 10 ordered RGBA PNG layers. Sume returns flat images at data[].url; n counts images, not layers.
- Ray 3.2 API aspect ratios and resolutions, and Sume's lists
Luma lists six Ray 3.2 aspect ratios (9:16 to 21:9) at 540p, 720p and 1080p. Sume reports ratios per model in supported_aspect_ratios and rejects size.
- Luma Scenes Ray 3.2 or Seedance 2: pin a model or use auto
Luma says Ray 3.2 and Seedance 2 are not interchangeable. On Sume, pin a catalog model for a known clip length, or send sume/auto and let Sume pick.
Written by Sume