Edit Video in Sume: which models appear and why H3 Max does not
Sume's Edit Video mode lists Auto, Kling 3.0, Wan 3.0 and MiniMax H3 and needs a reference video. H3 Max and Grok Imagine are missing; the API differs.

Switch the Agents Videos panel to Edit Video and the model list shrinks to Auto, Kling 3.0, Wan 3.0 and MiniMax H3; MiniMax H3 Max and Grok Imagine disappear. Edit mode needs a reference video, and the end-frame tile is hidden while you are in it.
That is the panel. The API has its own rules for video input, and they are broader, so a missing panel entry does not always mean the model cannot take a video.
Which models does Edit mode list?
The panel filters its model list by mode. Create mode shows every entry; Edit mode keeps only entries flagged for edit. The code comment for H3 Max says there is no Edit mode because the provider route behind it has no reference-to-video endpoint.
| Model | Create | Edit |
|---|---|---|
| Auto | Yes | Yes |
| Kling 3.0 | Yes | Yes |
| Wan 3.0 | Yes | Yes |
| MiniMax H3 | Yes | Yes |
| MiniMax H3 Max | Yes | No |
| Grok Imagine | Yes | No |
What does the panel's hint say?
The Edit Video option carries the hint "Reference video required", and Create Video carries "Text / image to video". A third mode, Avatar, is a separate product with its own model, covered in the avatar videos docs. The default model on switching to Edit is the first one in the Edit list, which is Auto.
What does the API accept for video input?
Through the API, video references go in input_references where the model lists video_url in supported_input_references, and in the Video Router docs other rows are built around a source video. The video docs say video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, and that Gemini Omni Flash 1.1 exposes its video edit mode through the Video Router video_url field.
Two special rows take exactly one source video: higgsfield-genjutsu (Motion Transfer, 4 to 30 seconds at 480p or 720p, with 1 to 8 reference images) and h3-max-recast (swap people, 5 to 30 seconds at 768p or 1080p, with 1 to 4 person photos). Neither appears in the panel's Edit list.
- Want to edit a clip in the panel: Auto, Kling 3.0, Wan 3.0 or MiniMax H3.
- Want to swap a person:
h3-max-recastthrough the API. - Want to transfer motion:
higgsfield-genjutsuthrough the API, when its provider is configured. - Want a prompt-driven edit of a clip up to 10 seconds: Gemini Omni Flash 1.1 via
video_url.
How should I choose?
Decide what the edit is. A light restyle or camera change on a short clip suits a model with video references. A person swap is a different row altogether. Check the length first: the Gemini Omni Flash 1.1 row runs 3 to 10 seconds, the recast row 5 to 30.
Use the video reference input post to compare the API rows and the Gemini Omni video edit post for product swaps.
What should I check before a batch?
Reference videos must be public HTTPS URLs the provider can fetch, and the docs list an unreachable reference as a first thing to check when a generation fails.
Submit one short clip first, open the finished job, and compare what you asked for with what came back. Use an Idempotency-Key on each attempt, because a replay with the same key returns the original job instead of creating and billing a second one. Only then queue the rest.
Sume reserves provider list times 1.25 when a job is submitted, and the poll response's usage.cost is the billable amount. Treat that field, not a panel estimate, as the number to budget with.
What does Sume not do?
Sume does not offer an Edit mode for Grok Imagine and does not list H3 Max under Edit in the panel, even though the API reports video references for H3 Max. If a model is missing from the panel, check the catalog before assuming it cannot take a video.
Sources
Related posts
More in Media tools
- ElevenLabs sound effects API: 0.5-30 s, loop, and Sume's route
ElevenLabs POST /v1/sound-generation takes duration_seconds 0.5-30 and a loop flag. Sume lists no sound-effect endpoint; here is the music route and its limits.
- English to Korean captions: style follows the text, not language
Localising a video to Korean? Sume's caption style follows the wording, and a Latin style on Hangul text is a 400. Pick styles per locale before a batch.
- Extract audio from an AI avatar video as a WAV with one API call
Sume audio detach pulls the track of one hosted video into a WAV or MP3 for $0.01. The request, the 16 kHz mono option, the caps and the no-audio error.
- Extract audio from a video for transcription: 16 kHz mono wav
Detach the audio from a Sume-hosted video as 16 kHz mono wav, then send it to speech-to-text. Costs, caps and the errors to expect.
Written by Sume