Which AI Video Models Can Edit an Existing Clip? V2V Options
Video-to-video in October 2026: Gemini Omni edit, Luma Modify, Runway Aleph 2, Grok video edit and Sume's person-swap and motion-transfer routes.

Video-to-video means different things on different pages. Some models restyle a clip by instruction, some swap the person but keep the motion, and some only use a clip as a reference for a fresh generation. Here is what each vendor and Sume document on 2026-10-03.
Instruction-based edit
Google's Omni documentation describes conversational editing through the Interactions API, with the caveat that audio in reference videos is ignored and voice editing is unsupported. On Sume, gemini-omni-flash-1.1 exposes its edit mode through the Video Router video_url field. The edit path cannot be combined with image or reference fields.
Elsewhere: Luma lists Modify Video for Ray 3.2 with clips up to 20 seconds depending on frame rate; Runway's pricing page lists Aleph 2 at 28 credits with a 56-credit minimum; and xAI's guide lists video-edit input for grok-imagine-video-1.5. Sume's Grok row, however, is image-to-video only.
Person swap and motion transfer
Two Sume routes take a source clip and keep its movement. h3-max-recast takes a source video of 5 to 30 seconds, with no shot over 15 seconds, plus 1 to 4 person photos; it keeps motion, camera, cuts and soundtrack, and the prompt is optional. higgsfield-genjutsu takes one source video plus 1 to 8 reference images for 4 to 30 seconds and is listed only when its provider is configured.
| Route | Source clip | Extra inputs | Sume per second |
|---|---|---|---|
| gemini-omni-flash-1.1 edit | video_url | Prompt only | $0.0375 to $0.375 by resolution tier |
| h3-max-recast | 5 to 30 s | 1 to 4 person photos | $0.375 (768p), $0.5625 (1080p) |
| higgsfield-genjutsu | 4 to 30 s | 1 to 8 images | $0.3975 (480p), $0.85125 (720p) |
Choosing
If you want to change wording, lighting or a prop with a sentence, use the instruction edit. If the choreography is right and the person is wrong, Recast is built for that and is priced accordingly at $0.375 per second. If you want a reference-driven remake rather than an edit, a fresh generation with video references is usually cheaper.
Sume's docs do not publish a quality ranking between these routes, so run a 5-second test on your own source before buying the full 30. See which model fits the inputs you have.
Sources
Related posts
More in Models
- Which AI video models take 1080p on Sume, and which do not
Seedance, Kling, Wan and Omni accept 1080p on Sume; H3 Max refines to it from native 768p; H3, Grok and Genjutsu stop lower. Full matrix.
- Which AI video models take a reference audio file on Sume?
Seedance 2.x, Wan 3.0, MiniMax H3 and H3 Max accept a reference audio clip on Sume; Gemini Omni Flash, Kling, Grok, Recast and Genjutsu do not.
- Which of Sume's nine aspect ratios work on Reels, Shorts, TikTok ads
Sume lists nine video aspect ratios. Checked against Instagram's 1.91:1 to 9:16 Reel range, YouTube's square-or-vertical Shorts rule and TikTok's ad ratios.
- Which Sume image model for a photo edit: mask, references or pixels?
Pick between ChatGPT Image 2.5, Ideogram 4.5 and Nano Banana 2 for a photo edit on Sume by what the edit needs: a mask, many references, or untouched pixels.
Written by Sume