Edit a video with a prompt and a photo: Omni edit takes no references
Gemini Omni Flash 1.1's video edit on Sume takes a video_url and a prompt only. To put a specific person from a photo into a clip, use h3-max-recast instead.

You cannot hand Gemini Omni Flash 1.1's edit mode a photo on Sume. The edit takes video_url and a text prompt, and the docs state that video_url cannot be combined with image_url, end_image_url or reference_*_urls (Video Router docs, read 2026-10-03). So "replace the man with this person" is not expressible there. For a specific person from a photo, use h3-max-recast, which exists for exactly that.
The edit mode is still useful. It is the prompt-driven video-to-video option for changes you can describe in words.
What Omni edit accepts
gemini-omni-flash-1.1 is one catalog id that Sume routes by the shape of the request. Send only a prompt and you get text-to-video; send video_url and you get an edit. In edit mode resolution is optional and defaults to 720p, aspect_ratio is rejected, and duration is only a hint for the reserve estimate (default 8 seconds) because output follows the source clip. Native audio is always on and generate_audio: false is rejected.
| Property | gemini-omni-flash-1.1 edit | h3-max-recast |
|---|---|---|
| Inputs | video_url + prompt | video_url + 1 to 4 photos, prompt optional |
| Photos | None; no references with video_url | One per person |
| How you name the new person | In words | With a photo |
| Resolutions | 360p, 720p, 1080p, 4K | 768p, 1080p |
| Output length | Follows the source; family is 3 to 10 s | Source length, 5 to 30 s |
| List price per second | $0.03 / $0.10 / $0.15 / $0.30 by resolution | $0.30 at 768p, $0.45 at 1080p |
| Sume multiplier | x 1.25 | x 1.25 |
A prompt edit that fits Omni
Omni edit fits changes like a prop, a color or a background: the docs' own example is "Replace the bottle with an apple. Keep everything else the same." It does not fit changes that depend on a likeness you cannot describe. The request below shows the edit shape; notice there is no photo field to add.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-edit-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Replace the red jacket with a green one. Keep everything else the same.",
"video_url": "https://example.com/clip.mp4",
"resolution": "720p",
"mode": "async"
}'How to choose
Use the question "can I describe the change in a sentence?" If yes and the clip is up to about 10 seconds, Omni edit is the cheap option: at 720p the list is $0.10 per second, about $0.125 through Sume. If the change is who the person is, use Recast and pay $0.375 per second at 768p. If you try to mix the two by sending photos with video_url to Omni, expect a rejection, not a degraded result.
One more difference is easy to miss. Omni's family is 3 to 10 seconds, while Recast accepts 5 to 30 seconds with no shot over 15. A 20-second clip can be recast but not edited by Omni in one request. For a pricing comparison across these two see Omni video edit reserve hint.
Where Omni edit still wins
It would be a mistake to read this as Omni being the weaker model. For a prompt-describable change on a short clip it is the better fit on price and flexibility: 4K output is available, the edit follows the source's length, and nothing needs a photo. Swapping a product, changing a garment color, replacing a background object or removing something are all things you can say in a sentence.
The line to hold is between changing what is in the scene and changing who someone is. The first is language; the second is a face. A prompt like "make the woman look like a famous actress" is both unreliable and not something to build a product on, and a prompt cannot reproduce a specific private person's likeness anyway. When identity is the point, a photo is the right input, and it should be a photo of someone who has agreed to appear.
Finally, remember the reserve behavior. Because Omni edit sends no duration to the provider, Sume reserves against a default of 8 seconds unless you pass a duration hint. If your source is 10 seconds, pass the hint so the hold matches the work.
Two honest limits
First, Sume does not expose a way to attach identity to an Omni edit. If your use case needs a particular real person's likeness, make sure you have that person's permission and use the model built for photo references. Second, sume/auto defaults to Gemini Omni Flash 1.1 for edit routing but never routes to h3-max-recast, so a person swap needs the explicit id on every call.
Sources
Related posts
More in Models
- End-frame AI video: which Sume models take a last frame
Seedance, Wan 3.0, Kling 3, MiniMax and Gemini Omni Flash accept a first and last frame; Grok Imagine and the swap rows do not. How to send both frames.
- Fastest Seedance or Kling model on Sume: latency labels
Sume labels seedance-2-fast and seedance-2-mini fast, seedance-2.5 and seedance-2 medium, and kling-3 medium-slow. What the labels mean and how to test.
- FLUX.2 pro prompt upsampling is on by default; Sume has no switch
fal's FLUX.2 pro page says automatic prompt upsampling is on by default. Sume exposes no passthrough field for it, so test exact-text prompts with n=4.
- 4:5 on FLUX.2 pro and Qwen Image becomes 1229x1536 on Sume
On Sume, FLUX.2 pro, Qwen Image and Recraft V4 turn custom ratios like 4:5 and 21:9 into pixels with a 1536 long side. The sizes table, and a 1080x1350 resize.
Written by Sume