MiniMax says H3 edits across modalities. Which Sume row edits a clip?
MiniMax lists reference and editing across modalities for H3. On Sume, H3 does text, image and reference video; clip edit by prompt is Omni's video_url mode.

MiniMax's H3 post lists reference and editing across image, video and audio inputs, but Sume's hosted minimax-h3 and minimax-h3-max rows do text-to-video, image-to-video (start and end frame) and reference-to-video, with no edit mode. To edit an existing clip with a prompt on Sume, send it as video_url to gemini-omni-flash-1.1.
Read on 2026-10-07: the MiniMax H3 post and Sume's Video Router docs.
Three rows accept video_url
For all three, video_url cannot be combined with start or end frames. The H3 rows do not accept it.
| Row | What video_url does | Prompt |
|---|---|---|
| gemini-omni-flash-1.1 | Edit source: change the clip as the prompt describes | Required |
| h3-max-recast | Person swap: 1 to 4 person photos replace the people | Optional, up to 2000 characters |
| higgsfield-genjutsu | Motion transfer source, with 1 to 8 reference images | Per its docs |
What H3 does on Sume
- Text-to-video, and image-to-video with a first and optional last frame.
- Reference-to-video with up to 9 images, 3 videos and 3 audio files.
- Video references are guidance for a new clip, not an edit of the source.
A rule for choosing
If the source clip must stay (same shot, same people, one change), use Omni's edit mode or Recast for a person swap. If you want a new clip that borrows the look of a reference clip, use H3's reference_video_urls. If you want to run MiniMax's own editing on the open weights, that is outside Sume.
| Task | Route |
|---|---|
| Change one object in a clip | Omni edit with video_url |
| Swap the person | h3-max-recast |
| New clip in the style of a reference | minimax-h3 with reference_video_urls |
Writing an edit prompt that holds
An edit prompt should name what changes and say what stays. For example, 'replace the red mug with a blue one; keep the hand, the table and the camera move.' Long, descriptive prompts invite the model to redo the whole shot, which is the opposite of an edit.
Send the source as video_url, keep the prompt short, and check the result against the source frame by frame. If a person must change, use the Recast route; if nothing in the source must survive, a new reference-to-video clip on H3 is the cleaner choice.
- One change per edit.
- State what stays.
- Compare first, middle and last frames.
Sources
Related posts
More in Models
- Minimum clip length by Sume video model: 2, 3, 4 or 5 seconds
Wan 3.0 starts at 2 s, Gemini Omni Flash 1.1 at 3 s, Seedance and Genjutsu at 4 s, MiniMax H3 and Recast at 5 s. Full min and max table for a Sora port.
- Music Router image_url: public HTTPS only, and null to clear it
Sume music image_url must be a public HTTPS image, and null clears it on a reused request object. When a mood still helps and when the prompt does more.
- Nano Banana 2.1 has no 512px size in Google's guide; Sume lists 512
Google's guide says 512px is unsupported on Nano Banana 2.1, yet Sume's catalog enum for it starts at 512. How to test what a call returns.
- Nano Banana 2.1 tracks 4 characters, 10 objects: Sume's cap
Google says Nano Banana 2.1 tracks 4 characters and 10 objects across up to 14 references. Sume's Image API caps input_references at 10, or 16 on GPT.
Written by Sume