Check MiniMax H3 reference limits before you submit: a Python guard
Sume's minimax-h3 rows take up to 9 images, 3 videos and 3 audios, 12 in all, and audio cannot be the only reference. A 20-line Python guard catches it.

A reference-to-video request to minimax-h3 or minimax-h3-max is valid on Sume when it has at most 9 images, 3 videos and 3 audio files, no more than 12 in total, and at least one image or video next to any audio. Each video and each audio clip must run 2 to 15 seconds, and the videos together, like the audios together, at most 15 seconds.
These limits are in the catalog row for each model (see the Video Router docs for the catalog route, read 2026-10-07). MiniMax's own release note describes a Ref2VA checkpoint for reference-to-video, which is the open-weights counterpart of that mode.
The rules in one table
On minimax-h3, Sume's catalog also notes that the first five reference images are free of an extra charge and each further image adds a list fee, so a guard that counts images also helps with cost.
| Input | Count | Length rule |
|---|---|---|
| Images | Up to 9 | None stated |
| Videos | Up to 3 | 2 to 15 s each, 15 s combined |
| Audios | Up to 3 | 2 to 15 s each, 15 s combined |
| All together | Up to 12 | Audio cannot be the only reference |
A guard you can run
The function takes the lists of lengths you already know and returns a list of problems. It makes no network call.
def check_h3_refs(n_images, video_secs, audio_secs):
errs = []
if n_images > 9:
errs.append("max 9 images")
for name, secs in (("video", video_secs), ("audio", audio_secs)):
if len(secs) > 3:
errs.append(f"max 3 {name} files")
if any(s < 2 or s > 15 for s in secs):
errs.append(f"each {name} must be 2-15 s")
if sum(secs) > 15:
errs.append(f"{name} total must be <= 15 s")
total = n_images + len(video_secs) + len(audio_secs)
if total > 12:
errs.append("max 12 references in all")
if audio_secs and not n_images and not video_secs:
errs.append("audio cannot be the only reference")
return errs
if __name__ == "__main__":
print(check_h3_refs(10, [4, 4, 8], [20]))
print(check_h3_refs(0, [], [6]))What the examples print
The first call flags the tenth image, the 14 references in all, the combined video length of 16 seconds and the 20-second audio. The second flags an audio-only request. Run it before you spend a reservation; the catalog lists these as constraints, and finding a break locally saves a round trip.
Where to put the check
- In the step that assembles the request body, not in a UI form.
- Beside the code that reads clip lengths, since you need them for the sums.
- Keep the numbers in one constant, and re-read the catalog when a model changes.
What the guard does not cover
The guard checks counts and lengths, which are the rules you can know without calling the server. It does not check file formats, whether a URL is reachable or whether a person in a photo is clear enough to use. Those failures come back from the job itself, so read the job's error message and fix the input, rather than adding more rules to the guard. Treat the guard as the cheap first filter and the server as the final word, and keep the two in step by re-reading the catalog whenever you see a new error.
Sources
Related posts
More in Developers
- Check an Omni edit's length with video-inspect before a timeline join
An edit should follow the source length. Confirm it with a probe-only video-inspect call before the clip goes into a Timeline render, with a Python read.
- Check duration, resolution, ratio against /v1/videos/models in Node
A Sora-era request will not fit every Sume model. A Node script reads GET /v1/videos/models and lists what the model rejects before you pay for a job.
- Check video model limits before you submit: GET /v1/videos/models
Models differ on length, resolution and references. Read GET /v1/videos/models and validate a request before a paid submit, with a Python preflight.
- Check GET /v1/videos/models before you pay: a Python validator
Read GET /v1/videos/models and reject bad duration, resolution or audio flags in Python before POST /v1/videos. Fewer 400s, no paid retries.
Written by Sume