How to remove an object from a video with AI
Name the object in a video-to-video AI edit prompt and say what stays. How to do it on Sume, how to swap or add objects, or blur them instead.

To remove an object from a video with AI, send the clip to a video-to-video (edit) model with a prompt that names the object and says what stays, for example “Remove the coffee cup from the table. Keep everything else the same.” The model rewrites the clip from that instruction, so you draw no mask and track nothing by hand. Nothing guarantees a clean result, though: check the spot where the object was.
On Sume, that is the edit mode of gemini-omni-flash-1.1, and the docs' own edit example is a swap of this kind: “Replace the bottle with an apple. Keep everything else the same.” If the object only has to be hidden, blurring or boxing a fixed area is the non-generative option. Facts come from the Video Router and Video filter docs, the Sume API reference, and Sume's API code, read on 2026-09-28.
How do I remove an object from a video with Sume?
Send the clip's public HTTPS URL as video_url to POST /v1/video-router/generate with model: "gemini-omni-flash-1.1", currently the only model with this edit, and put the removal in prompt. Edit a video with a prompt covers the other fields, why the edit uses this route, and how to fetch the result.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: remove-cup-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Remove the red coffee cup from the left side of the table. Keep everything else the same.",
"video_url": "https://example.com/desk-clip.mp4",
"resolution": "720p",
"mode": "async"
}'How should I word the prompt?
Name the object so it can't be mistaken for anything else in the frame: what it is, its color, and where it sits. Then say what stays. The same pattern covers asking to replace an object or add one, since the prompt describes the edit. Examples to adapt:
- A stray object: “Remove the red coffee cup from the left side of the desk. Keep everything else the same.”
- A person in the background: “Remove the man in the blue jacket walking behind her. Keep her and the street the same.”
- A swap: “Replace the soda can with a glass of water. Keep the hand and the table the same.”
- An addition: “Add a small potted plant on the right side of the desk. Keep everything else the same.”
Will the spot where it was look right?
Not guaranteed. The API reference describes the edit as a source clip “rewritten by prompt”, and the docs show a swap, not a removal, with no promise about how the gap is filled. If a removal leaves something odd, try naming what should be there instead, such as “show the bare wooden table where the cup was.”
The docs also don't give a maximum source length, or say whether the output keeps the source's length, framing, or original sound. They do say the model always generates audio, with no switch to turn it off. Watch the edited clip against the original before you use it.
Can I blur or cover the object instead?
Yes, when the object stays in one place and the clip is already your workspace's media.sume.com file, such as an earlier Sume job's output. Today a Video filter job can blur that fixed area or cover it with a solid box, without generating anything. How to blur part of a video has the filtergraph for both.
| AI edit: remove or replace | Video filter: blur or box | |
|---|---|---|
| Route | POST /v1/video-router/generate | POST /v1/video-filter |
| Source clip | A public HTTPS URL | Your workspace's media.sume.com clip |
| What changes | The whole clip, rewritten from your prompt | Only what the program changes; size, frame rate, and audio carry over |
| Sound | Always generated; keeping the original sound isn't documented | The source's audio |
| Source length | Not documented | Up to 300 seconds |
How much does it cost?
An edit is billed per second of output, by resolution, at the provider's list price × 1.25: $0.125 a second at the default 720p, plus a 5.5% agent fee by default. A filter job is billed per encode, and its /check call is free; the docs say to confirm the rate in GET /v1/catalog.
Sources
Related posts
More in Models
- Runway Act-Two: what it does, API inputs and cost
Runway Act-Two applies an acted performance from a 3–30 second video to a character image or video. API fields, credit cost, and a Sume option.
- Runway Aleph 2.0 API: aleph2 inputs, limits and price
Yes: Aleph 2.0 is aleph2 on Runway's POST /v1/video_to_video. It edits a 2–30 s video from a prompt and keyframes at 28 credits a second.
- Runway API Seedance: calling Seedance 2.5 and 2.0 on Runway
Yes, Runway's API serves Seedance: seedance2_5, seedance2, seedance2_fast and seedance2_mini. Inputs, limits and credits, and other ways to call Seedance.
- Runway Gen-4.5 vs Seedance 2.0: inputs, length and price
On Runway's API, Gen-4.5 makes 2–10 s clips for 12 credits a second; Seedance 2.0 makes 4–15 s clips with audio for 36–150 credits a second.
Written by Sume