Add sunglasses or a hat to a person in an existing video (Omni edit)
Use Gemini Omni Flash 1.1 video_url edit on Sume to add an accessory to a person already in a clip. Prompt wording, a frame check list, and 4-second prices.

To add sunglasses or a hat to a person in an existing video on Sume, send the clip as video_url to gemini-omni-flash-1.1 with a prompt that names the accessory and says what must not change. The edit route keeps the source clip as its base, so you describe the addition and the keep clause, not a full scene. Expect to check the result frame by frame, because accessories sit on the face and head, where motion and occlusion are hardest.
Read this as a recipe and not as a guarantee. The Sume docs describe the route, its limits and its price. They do not promise that any one accessory will stay stable across a clip. That is the part you verify after each run.
A prompt that names the accessory and the keep clause
Be specific about the object and its color, because an edit has no reference image to copy from. Edit mode takes only a prompt and the source clip: the Video Router doc says a video_url request cannot be combined with image_url, end_image_url or reference_*_urls. So the look of the sunglasses or the hat has to come from words.
- Add round black sunglasses to the man in the blue shirt. Keep his face, his movement and the background the same.
- Put a plain grey wool beanie on the woman. Keep her hair color, her expression and the camera movement the same.
- Add a red baseball cap with no logo. Keep everything else the same.
{
"model": "gemini-omni-flash-1.1",
"prompt": "Add round black sunglasses to the man in the blue shirt. Keep his face, his movement and the background the same.",
"video_url": "https://example.com/source.mp4",
"resolution": "720p",
"duration": 4
}Frames to check before you accept a take
The risk with a face-adjacent edit is not the first frame. It is the moments when the head turns, the hand crosses the face, or the person walks out of the light. Pull still frames from a few moments and compare them to the source.
| Moment | What to look at | If it fails |
|---|---|---|
| First second | Is the accessory present and the right color | Name the color and shape again in the prompt |
| Head turn | Does it stay attached and keep its shape | Re-run once, since results differ each run |
| Hand near face | Does it appear in front of or behind the hand correctly | Pick a source clip with less occlusion |
| Last second | Has it drifted or changed style | Shorten the clip, or edit the last part separately |
| Audio | Does the sound still match the picture | Edit mode always outputs native synced audio, so listen to it |
Why a reference photo is not an option here
If the hat must match a real product, such as a brand cap, the edit route cannot take a product photo. Reference-to-video exists on the same model id, but that is a new generation: image_url and reference_image_urls select a different mode. The same question comes up in Omni edit with a photo, where the answer is the same, and the sibling page on image reference tags shows how the reference mode names its inputs.
For the edit mode itself, remember what the route fixes: no aspect_ratio (a 400 if sent), no duration sent to the provider, and generate_audio: false is a 400.
Price of a 4 second pass
The catalog lists the provider rates per second at $0.03 for 360p, $0.10 for 720p, $0.15 for 1080p and $0.30 for 4K, and Sume bills those rates times 1.25. The edit reserve follows the duration you send as a hint (default 8 s), so send the real length. For a 4 second clip the arithmetic is below, rounded up to the cent.
A short pass is cheap enough to run twice. If the first take loses the accessory in the head turn, the second costs the same again, so budget for two takes per clip when the person moves a lot. More on that trade-off in fixing one detail in a finished clip.
Resolution is a request field, not a prompt word. The default is 720p. If the clip is going into a feed at 1080p, ask for 1080p in the request and expect the higher row. If you only want to learn whether the prompt works, a 360p test is the cheapest way to see it, and then you spend the 720p or 1080p money once.
| Resolution | Provider list per second | Sume per second | 4 second pass |
|---|---|---|---|
| 360p | $0.03 | $0.0375 | $0.15 |
| 720p | $0.10 | $0.125 | $0.50 |
| 1080p | $0.15 | $0.1875 | $0.75 |
| 4K | $0.30 | $0.375 | $1.50 |
Sources
Related posts
More in Use cases
- AI play-by-play voiceover for a sports highlight reel: how to time it
Write one line per play, ask TTS for sentence timings, and place each clip on its line: a 60-second reel costs about 14 cents on Sume. Not live commentary.
- AI spokesperson release notes, 90 seconds: two jobs, cost by tier
A 90-second spokesperson video is two Sume avatar jobs because one job tops out at 60 seconds. Cost at standard, plus and max, and how to split the script.
- How do I make a spoken wake-up alarm with an AI voice and music?
A wake-up alarm is a short voice greeting joined to a music file: 2 cents of TTS and a 1-cent concat per day, with one $0.125 music track reused all week.
- Animate a painting with AI video: first frame and a slow camera prompt
Pin a painting as the first frame with frame_images on Sume, ask for a slow camera move, and pick the model by length and price. A 6-second price table.
Written by Sume