Edit a video with a prompt and a photo: Omni edit takes no references

Gemini Omni Flash 1.1's video edit on Sume takes a video_url and a prompt only. To put a specific person from a photo into a clip, use h3-max-recast instead.

6 min readSume
All posts

You cannot hand Gemini Omni Flash 1.1's edit mode a photo on Sume. The edit takes video_url and a text prompt, and the docs state that video_url cannot be combined with image_url, end_image_url or reference_*_urls (Video Router docs, read 2026-10-03). So "replace the man with this person" is not expressible there. For a specific person from a photo, use h3-max-recast, which exists for exactly that.

The edit mode is still useful. It is the prompt-driven video-to-video option for changes you can describe in words.

What Omni edit accepts

gemini-omni-flash-1.1 is one catalog id that Sume routes by the shape of the request. Send only a prompt and you get text-to-video; send video_url and you get an edit. In edit mode resolution is optional and defaults to 720p, aspect_ratio is rejected, and duration is only a hint for the reserve estimate (default 8 seconds) because output follows the source clip. Native audio is always on and generate_audio: false is rejected.

Omni edit vs Recast, read 2026-10-03
Propertygemini-omni-flash-1.1 edith3-max-recast
Inputsvideo_url + promptvideo_url + 1 to 4 photos, prompt optional
PhotosNone; no references with video_urlOne per person
How you name the new personIn wordsWith a photo
Resolutions360p, 720p, 1080p, 4K768p, 1080p
Output lengthFollows the source; family is 3 to 10 sSource length, 5 to 30 s
List price per second$0.03 / $0.10 / $0.15 / $0.30 by resolution$0.30 at 768p, $0.45 at 1080p
Sume multiplierx 1.25x 1.25

A prompt edit that fits Omni

Omni edit fits changes like a prop, a color or a background: the docs' own example is "Replace the bottle with an apple. Keep everything else the same." It does not fit changes that depend on a likeness you cannot describe. The request below shows the edit shape; notice there is no photo field to add.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-edit-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Replace the red jacket with a green one. Keep everything else the same.",
    "video_url": "https://example.com/clip.mp4",
    "resolution": "720p",
    "mode": "async"
  }'

How to choose

Use the question "can I describe the change in a sentence?" If yes and the clip is up to about 10 seconds, Omni edit is the cheap option: at 720p the list is $0.10 per second, about $0.125 through Sume. If the change is who the person is, use Recast and pay $0.375 per second at 768p. If you try to mix the two by sending photos with video_url to Omni, expect a rejection, not a degraded result.

One more difference is easy to miss. Omni's family is 3 to 10 seconds, while Recast accepts 5 to 30 seconds with no shot over 15. A 20-second clip can be recast but not edited by Omni in one request. For a pricing comparison across these two see Omni video edit reserve hint.

Where Omni edit still wins

It would be a mistake to read this as Omni being the weaker model. For a prompt-describable change on a short clip it is the better fit on price and flexibility: 4K output is available, the edit follows the source's length, and nothing needs a photo. Swapping a product, changing a garment color, replacing a background object or removing something are all things you can say in a sentence.

The line to hold is between changing what is in the scene and changing who someone is. The first is language; the second is a face. A prompt like "make the woman look like a famous actress" is both unreliable and not something to build a product on, and a prompt cannot reproduce a specific private person's likeness anyway. When identity is the point, a photo is the right input, and it should be a photo of someone who has agreed to appear.

Finally, remember the reserve behavior. Because Omni edit sends no duration to the provider, Sume reserves against a default of 8 seconds unless you pass a duration hint. If your source is 10 seconds, pass the hint so the hold matches the work.

Two honest limits

First, Sume does not expose a way to attach identity to an Omni edit. If your use case needs a particular real person's likeness, make sure you have that person's permission and use the model built for photo references. Second, sume/auto defaults to Gemini Omni Flash 1.1 for edit routing but never routes to h3-max-recast, so a person swap needs the explicit id on every call.

Sources

Related posts

More in Models

All Models posts

Written by Sume