Omni Flash vs H3 Max for reference-to-video: limits side by side

Gemini Omni Flash 1.1 takes 10 images and 3 short videos; MiniMax H3 Max adds audio refs, 12 files total. Limits, tags and a sample request on Sume.

5 min readSume
All posts

Gemini Omni Flash 1.1 and MiniMax H3 Max both do reference-to-video on Sume, but their limits differ in ways that decide which one you can use. Omni takes up to 10 reference images and up to 3 reference videos of at most 3 seconds each, and no audio references, for 3 to 10 second clips. H3 Max takes up to 9 images, 3 videos and 3 audio files, 12 files in total, for 5 to 15 second clips, and an audio file cannot be your only reference.

The numbers below come from Sume's Video Router docs and Video generation docs, read on 2026-10-03. Both models ship in the catalog under the ids gemini-omni-flash-1.1 and minimax-h3-max.

How do the two models compare on references?

Read the table by what you hold. If you have a character photo and a short reference clip, either model can take them. If you also have a voice sample, only H3 Max accepts it. If you have a single long reference clip, Omni's 3-second cap per video is the constraint, and you would trim the clip first.

Both models always produce audio of their own. Neither lets you switch it off: generate_audio: false is rejected on Omni, and H3 Max's audio is always on, so you omit the flag on both.

Reference-to-video limits on Sume, from Sume docs, read 2026-10-03
Limitgemini-omni-flash-1.1minimax-h3-max
Reference imagesUp to 10Up to 9
Reference videosUp to 3, each 3 s or lessUp to 3
Reference audioNoneUp to 3; cannot be the only reference
Total filesPer-type limits only12 across all types
Clip length3-10 s5-15 s
Resolutions360p, 720p, 1080p, 4K480p, 768p (default), 1080p (latent refinement)
Aspect ratio16:9 or 9:16Read supported_aspect_ratios from the catalog
Audio on outputAlways onAlways on

How are references addressed in the prompt?

Omni addresses references by position. Name a reference image as <IMAGE_REF_0> and a reference video as <VIDEO_REF_0>, counted from zero in list order, and write the prompt around those tags. The Sume docs example is a cat inspired by <IMAGE_REF_0> walking through the setting shown in <VIDEO_REF_0>.

H3 Max uses the same request field names, reference_image_urls, reference_video_urls and reference_audio_urls. Because Sume's docs do not describe a tag syntax for H3 Max, describe each reference in plain words and check the first result before you write a batch around it.

curl -X POST https://api.sume.com/v1/video-router/generate \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: omni-ref-001" \
  -d '{
    "model": "gemini-omni-flash-1.1",
    "prompt": "A cat inspired by <IMAGE_REF_0> walks through the setting shown in <VIDEO_REF_0>.",
    "reference_image_urls": ["https://example.com/inputs/cat.png"],
    "reference_video_urls": ["https://example.com/inputs/setting-3s.mp4"],
    "resolution": "720p",
    "duration": 8,
    "aspect_ratio": "16:9",
    "mode": "async"
  }'

Which should I pick for which job?

Pick on what the clip needs rather than on the model's reputation. The rules below follow the limits, not a quality ranking, and Sume's docs publish no head-to-head benchmark between the two.

  • Choose Omni when you need 4K or a 360p draft, when a clip must be 3 or 4 seconds, or when you also want to edit an existing video with video_url.
  • Choose H3 Max when you have a voice or sound reference, when you need a 12 to 15 second clip, or when you need up to 9 images plus videos in one request.
  • Choose neither if your reference is a single clip longer than 3 seconds and you cannot trim it; Seedance 2.5 and Wan 3.0 have their own video-reference rules.
  • Keep video_url and reference_video_urls apart on Omni: video_url is an edit source and cannot be combined with image_url, end_image_url or any reference_*_urls.

How are the two billed?

Both bill per output second by resolution, so the same 8-second clip costs different amounts depending on the resolution you pick. The Omni docs state provider list times 1.25 per output second by resolution. Do not carry a number from one model to the other: ask GET /v1/videos/models for each model's pricing_skus, and multiply by your clip length. A 360p Omni draft and a 4K Omni final are different rows of the same model, which is the cheapest way to test a reference set before you spend on the final.

Check the quoted amount on a submit with your real reference set before you run a batch, and budget for retakes: a reference set needs a few trial clips before it behaves.

What about audio references on other models?

Reference audio is honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max, according to the Video generation docs. Omni and higgsfield-genjutsu accept video references but not audio. The stored post which Sume models honor audio_url walks through that list, and Gemini Omni Flash 1.1 API: text, image, and reference video modes covers Omni's other modes.

One safety check applies to both: use public HTTPS URLs for every reference, and keep them reachable until the job finishes. Sume returns input_media_unreachable when it cannot fetch or mirror input media, so check each URL in a browser first.

Sources

Related posts

More in Models

All Models posts

Written by Sume