MiniMax says H3 edits across modalities. Which Sume row edits a clip?

MiniMax lists reference and editing across modalities for H3. On Sume, H3 does text, image and reference video; clip edit by prompt is Omni's video_url mode.

4 min readSume
All posts

MiniMax's H3 post lists reference and editing across image, video and audio inputs, but Sume's hosted minimax-h3 and minimax-h3-max rows do text-to-video, image-to-video (start and end frame) and reference-to-video, with no edit mode. To edit an existing clip with a prompt on Sume, send it as video_url to gemini-omni-flash-1.1.

Read on 2026-10-07: the MiniMax H3 post and Sume's Video Router docs.

Three rows accept video_url

For all three, video_url cannot be combined with start or end frames. The H3 rows do not accept it.

Rows that accept video_url on Sume (read 2026-10-07)
RowWhat video_url doesPrompt
gemini-omni-flash-1.1Edit source: change the clip as the prompt describesRequired
h3-max-recastPerson swap: 1 to 4 person photos replace the peopleOptional, up to 2000 characters
higgsfield-genjutsuMotion transfer source, with 1 to 8 reference imagesPer its docs

What H3 does on Sume

  • Text-to-video, and image-to-video with a first and optional last frame.
  • Reference-to-video with up to 9 images, 3 videos and 3 audio files.
  • Video references are guidance for a new clip, not an edit of the source.

A rule for choosing

If the source clip must stay (same shot, same people, one change), use Omni's edit mode or Recast for a person swap. If you want a new clip that borrows the look of a reference clip, use H3's reference_video_urls. If you want to run MiniMax's own editing on the open weights, that is outside Sume.

Task to Sume route (read 2026-10-07)
TaskRoute
Change one object in a clipOmni edit with video_url
Swap the personh3-max-recast
New clip in the style of a referenceminimax-h3 with reference_video_urls

Writing an edit prompt that holds

An edit prompt should name what changes and say what stays. For example, 'replace the red mug with a blue one; keep the hand, the table and the camera move.' Long, descriptive prompts invite the model to redo the whole shot, which is the opposite of an edit.

Send the source as video_url, keep the prompt short, and check the result against the source frame by frame. If a person must change, use the Recast route; if nothing in the source must survive, a new reference-to-video clip on H3 is the cleaner choice.

  • One change per edit.
  • State what stays.
  • Compare first, middle and last frames.

Sources

Related posts

More in Models

All Models posts

Written by Sume