Video-to-video clip length limits on Sume, by tool
Recast takes 5 to 30 s, Genjutsu 4 to 30 s, Omni edit 3 to 10 s, Avatar face swap about 4 to 15 s. A table and a Python picker choose by clip length.

The source-clip length you have decides which Sume video-to-video tool you can use. h3-max-recast takes 5 to 30 seconds with no shot over 15, higgsfield-genjutsu takes 4 to 30 seconds, gemini-omni-flash-1.1 edit works within a 3 to 10 second family, and the Beta Avatar Face Swap targets about 4 to 15 seconds with usable audio (Video Router docs and Face swap docs, read 2026-10-03). A 3-second clip fits only Omni; a 45-second clip fits none of them until you cut it.
Length is only the first filter, but it is the one that fails fastest and most plainly, so check it before anything else.
The limits in one table
Each row restates what the docs say about the source video, not about generation from text. Genjutsu's duration is the input video's length rounded up, and Recast's is the source length rounded up as well. Omni's edit sends no duration to the provider, so its output follows the source and the family bound is the practical limit.
| Tool | What it does | Source length | Other limits |
|---|---|---|---|
h3-max-recast | Swaps people with 1 to 4 photos | 5 to 30 s | No shot over 15 s; 768p, 1080p |
higgsfield-genjutsu | Motion Transfer with 1 to 8 images | 4 to 30 s | 480p, 720p; listed only when configured |
gemini-omni-flash-1.1 edit | Prompt-driven edit | 3 to 10 s family | No references with video_url; 360p to 4K |
| Avatar Face Swap (Beta) | Applies a ready avatar's face | about 4 to 15 s planned | Needs usable audio; quality required |
seedance-2.5 reference | Guides a new clip | 4 to 30 s generated | Not an edit of your clip |
A picker you can run
This function takes the clip length and the longest single shot and returns the tools that accept it. It encodes only the numbers in the table above, so update it when the docs change.
def options(seconds: float, longest_shot: float | None = None) -> list[str]:
out = []
shot_ok = longest_shot is None or longest_shot <= 15
if 5 <= seconds <= 30 and shot_ok:
out.append("h3-max-recast")
if 4 <= seconds <= 30:
out.append("higgsfield-genjutsu")
if 3 <= seconds <= 10:
out.append("gemini-omni-flash-1.1 (edit)")
if 4 <= seconds <= 15:
out.append("avatar face swap (beta)")
return out
for clip in (3, 4.5, 12, 20, 45):
print(clip, options(clip, longest_shot=clip))What the outputs show
Running it, the 3-second clip returns only Omni edit. The 4.5-second clip returns Genjutsu, Omni and face swap, but not Recast, because Recast needs 5 seconds. The 12-second clip returns all four. The 20-second clip, taken as one unbroken shot, returns only Genjutsu, since Recast rejects a shot over 15 seconds and the others cap sooner; if it were cut into shots under 15 seconds, Recast would qualify too. The 45-second clip returns nothing, which is the signal to trim or split.
These bounds describe acceptance, not suitability. Each tool does a different job. Recast changes who is in the footage, Genjutsu transfers motion onto images, Omni edits by prompt and face swap applies one avatar's face. See Recast vs Genjutsu for the two person-swap options in detail.
Three clips, three decisions
Real sources do not arrive at tidy lengths, so here are three worked cases. A 4.2-second reaction clip from a creator is too short for Recast and fits Genjutsu or Omni, so the question becomes whether you need to replace the person (Genjutsu moves your images with the source's motion) or edit the scene (Omni, by prompt). A 22-second product demo with two cuts, each shot under 15 seconds, fits Recast and Genjutsu; choose Recast if the goal is the same footage with a different person, since it keeps cuts and sound. A 40-second interview fits nothing as one piece: pick the 12 to 20 seconds that matter, trim them, and process that.
The cost difference follows the same logic. Recast at 768p is $0.375 per second through Sume, so the 22-second demo is about $8.25. Omni edit at 720p is $0.125 per second through Sume, so a 10-second clip is about $1.25 but cannot take a photo. The cheaper tool is only cheaper if it does the job.
What to do outside the window
Shorter than the minimum: you cannot pad a source with silence and expect the model to like it. Re-capture a slightly longer take, or use a tool with a lower floor. Longer than the maximum: cut with video trim into ranges and process each range, accepting that you will need to join the results afterward. A shot over 15 seconds that Recast rejects can be split at a natural pause, which also gives the model two simpler jobs.
- Probe every source before quoting a cost.
- Round lengths up; billing does.
- Record the tool and its limits in your job log so a later maintainer can see why a clip went where it did.
Sources
Related posts
More in Models
- Vietnamese, Thai, Indonesian, Malay TTS API: vi, th, id, ms on Sume
Cartesia Sonic 3.6 lists vi, th, id and ms. How to send each as the language on Sume TTS, why the field is required, and a per-character cost for each script.
- What happens to the audio in each Sume video-to-video tool
Recast and Avatar Face Swap keep the source audio, Kling motion control keeps it by default, Omni always produces native audio, Genjutsu takes no audio field.
- Where to try MAI-Voice-2.1 before paying: a 10-minute listening test
Microsoft lists the MAI Playground, Copilot Audio Expressions and Foundry for MAI-Voice-2.1. Run a ten-minute listening test with a fixed script and cost it.
- Which AI video models accept an input video on Sume?
Seedance, Wan 3.0, H3, H3 Max and Gemini Omni Flash take video references; Recast, Genjutsu and Omni edit need a source video. Kling and Grok take none.
Written by Sume