Flow edit modes Insert, Remove, Camera: the prompt wording for Sume

Flow's Omni edit offers Insert, Remove and Camera modes; Sume has one video_url edit and a prompt. Prompt templates for each mode and what to check.

5 min readSume
All posts

Flow's edit and refine screen lists Insert, Remove and Camera among its edit modes. Sume has no mode switch: you send one video_url and a prompt to gemini-omni-flash-1.1, and the wording of the prompt is the mode. This page gives a template for each of the three modes and says where the result needs a human check.

What does Flow list, and what does Sume take?

Google's Flow help page names the edit modes Insert, Remove and Camera, and says edit and refine require Gemini Omni Flash and select up to a 10-second segment. The page I read gives only those three names verbatim; do not assume a longer list.

On Sume the Video Router doc describes edit mode as video_url plus a prompt that describes the edit, with an optional resolution and no aspect_ratio or duration. The source cannot be combined with image_url, end_image_url or the reference fields, so a reference photo of the object to insert is not an option in the same call.

What prompt do I write for each mode?

Name what changes, name what stays. Omni's page says conversational editing carries context across turns, but Sume sends each edit fresh, so repeat the anchor every time.

The templates below are starting points, not documented syntax. Test them on your own footage.

Edit templates for gemini-omni-flash-1.1, read 2026-10-03
Flow modePrompt templateCheck afterwards
InsertAdd a red bicycle leaning on the wall at the right. Keep everything else the same.The object stays put across frames
RemoveRemove the person in the grey jacket. Fill the space naturally. Keep everything else the same.No ghosting where the person was
CameraKeep the scene and action. Change to a slow dolly-in from the same position.The subject and lighting are unchanged

What does the request look like?

The only thing that changes between modes is the prompt string. mode: async returns a job to poll, and a retry with the same Idempotency-Key returns the original job.

If the clip is longer than 10 seconds, cut it first; Google's API page caps edit inputs at 10 seconds.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-remove-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "Remove the person in the grey jacket. Keep everything else the same.",
    "video_url": "https://media.sume.com/artifacts/artf_demo/street.mp4",
    "resolution": "1080p",
    "mode": "async"
  }'

How do I check the result?

Use video inspect to pull stills at the moments that matter, such as the first and last second, and compare them with the source. An edit that changed more than you asked for is the common failure, and a prompt that restates what must stay is the usual fix.

Sume does not mask a region or draw a selection box, so an edit that must touch only one corner of the frame depends entirely on the prompt. When exact placement matters, Flow's visual selection is the better tool.

What about text overlays?

Flow's page I read lists text overlays among its editing capabilities, but I do not have wording for them beyond that line. Sume's Video Router doc makes no promise about rendering exact text, so for text that must be spelled exactly, add it in a step you control rather than asking the model to draw it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume