reference_video_urls or video_url? Reference footage vs edit source
On Sume, reference_video_urls guide a new clip; video_url is a source you edit or swap. They cannot be combined on Gemini Omni Flash. Which model takes which.

Use reference_video_urls when you want a new clip guided by existing footage, and video_url when you want to change the footage itself. On Sume the two do different jobs: references go to Seedance, Wan, MiniMax and Gemini Omni Flash for reference-to-video, while video_url is the source for Gemini Omni Flash edits, Genjutsu Motion Transfer and H3 Max Recast.
What each field means
A reference clip adds information to a new generation. An edit source is the thing being edited, so its length and framing carry through.
| Model id | reference_video_urls | video_url |
|---|---|---|
| wan-3.0, minimax-h3, minimax-h3-max | yes | no |
| seedance-2.5 and Seedance 2.0 rows | yes | no (no video edit) |
| gemini-omni-flash-1.1 | yes, up to 3 clips of 3 s | yes, an edit source |
| higgsfield-genjutsu | no | yes, plus 1 to 8 reference images |
| h3-max-recast | no | yes, plus 1 to 4 photos |
| kling-3, grok-imagine-video-1.5 | no | no |
Gemini Omni Flash rules
The docs are explicit: video_url is the edit source, not a reference, and it cannot be combined with image_url, end_image_url or reference_*_urls. An edit sends a prompt describing the change and optionally a resolution, defaulting to 720p; it takes no aspect ratio or duration.
Choosing
To keep a performance and change the person, use Recast. To move your images with a clip's motion, use Genjutsu. To change one thing in a clip, use a Gemini Omni Flash edit. To make a new clip that borrows style or motion from footage, use reference video on Wan, MiniMax or Seedance.
Sources
Related posts
More in Developers
- Sora video ids in your database after the shutdown: what to keep
OpenAI's Videos API shut down 2026-09-24. Old video_ids no longer resolve, so store your own file URL and the model used. Schema fields included.
- Speech-to-text audio too large? Sume STT takes up to 10 MB, hosted
Sume's stt_create needs a public HTTPS audio URL on the Sume media host, 10 MB at most. What fits: 16 kHz mono wav versus mp3, and how to cut a long file.
- SSML in text to speech: Sume takes a plain transcript, no ssml field
Does Sume's text to speech accept SSML? The tts_create body has a plain transcript and rejects unknown keys. What to use for speed, volume, emotion and pauses.
- Sume API rate limits by plan: requests per minute for writes and reads
Sume gives every API key a per-minute budget set by plan: 120 writes on Free up to 1200 on Scale, with reads at forty times the write number. Table and headers.
Written by Sume