Remove background from video with AI: what Sume covers
Sume has no video background-removal endpoint: RMBG 1.0 takes a still image. What you can do instead is a prompt edit with video_to_video, or a crop.

Sume does not ship a video background-removal or matting endpoint. Sume RMBG 1.0 takes a public HTTPS image_url, so it works on stills only. For video, the closest options are a prompt edit through video_to_video and a crop. Adobe's September 2026 Firefly notes describe a Remove Background for video clips that works across the whole clip without frame-by-frame masking; Sume has no equivalent.
Sume facts are from Video Router, Video frames and Video filter, read 2026-09-30.
What does RMBG 1.0 take?
The OpenAPI description reads: "Sume RMBG 1.0 background-removal request. Provide a public HTTPS image_url." There is no video_url field. You can pull a still from a clip with video frames, remove its background, and use the result as an image, but that gives you one cut-out, not a matted video.
Can a prompt change the background of a clip?
Yes, as an edit, not a matte. The gemini-omni-flash-1.1 model on Video Router accepts video_url as the edit source, and the prompt describes the edit. The docs' own example is "Replace the bottle with an apple. Keep everything else the same." You would describe the background change the same way and review the result. The output is a new generated video, so it can differ from the source in ways a mask would not.
| Field | Rule |
|---|---|
video_url | The edit source; cannot combine with image_url, end_image_url or reference_*_urls |
prompt | Describes the edit |
resolution | Optional, default 720p |
aspect_ratio / duration | Not accepted on this capability |
What about cropping the background out?
If the background is a border, video filter's crop op takes a rectangle as fractions of the source frame (x, y, width, height). That removes pixels outside the rectangle; it cannot cut around a moving subject.
What should I do for a transparent or masked result?
Sume does not return a transparent video. If you need a matte, do that step in an editor, then bring the result back as a Sume-hosted clip for trimming or assembly. For stills, see remove backgrounds from images in bulk. Frame extraction is capped at 24 frames per call, so it is a way to get references, not a per-frame matting pipeline.
Sources
Related posts
More in Use cases
- Replace video audio with an AI voice (and the lip-sync catch)
Sume has no one-call audio swap: detach the audio, transcribe it, make a new voice with TTS, and lay it on a Timeline. Lips will not re-sync to the new voice.
- AI ad resizer: one image, several aspect ratios, one API
Sume has no resizer button. Send the same source image to POST /v1/images once per aspect_ratio (1:1, 4:5, 9:16, 16:9) and recompose each ad slot.
- Seedance 2.5 draft mode: a 480p-then-1080p loop on Sume
Higgsfield added Seedance 2.5 Draft Mode. On Sume there is no draft model: render at resolution 480p, pick a take, then re-request at 1080p.
- Seedream 5.0 pro bbox and point tags vs region edits on Sume
Seedream 5.0 pro edits a region from <bbox> and <point> tags on a 0-999 grid. Sume lists other Seedream ids; use mask_url on ChatGPT Image 2.5 for regions.
Written by Sume