MiniMax H3-Context-IR returns a prompt, not a video: Sume view

MiniMax's H3-Context-IR endpoint reads text, images, audio and video and returns an enhanced prompt without making a video. What Sume has instead.

5 min readSume
All posts

H3-Context-IR is a MiniMax endpoint, POST /v2/h3_context_ir, that reads your text, images, video and audio and returns an enhanced video prompt without creating a video. Sume's docs list no equivalent standalone prompt-enhancer endpoint, so on Sume you write the prompt yourself or ask an agent to write it, then submit minimax-h3 with the finished text.

Everything about the MiniMax endpoint below is from its page, Create H3-Context-IR Task, read on 2026-10-02. Sume's side is from Video generation and Video Router.

What does H3-Context-IR actually do?

According to the page, it interprets multimodal context across text, images, audio and video and generates an enhanced video prompt, and it does not create a video generation task. The request takes model (only MiniMax-H3), a content array, a required duration of 4 to 15 seconds, an optional ratio, and an optional callback_url. It returns a task_id with task type h3_context_ir; you fetch the finished prompt from content.prompt through the query-task endpoint.

It uses the same media limits as video generation: images up to 30 MB, videos up to 50 MB and 3 clips, audio up to 15 MB and 3 clips, and 64 MB for the whole request. The text item is limited to 7,000 characters.

H3-Context-IR request fields (read 2026-10-02)
FieldRequiredValue on the page
modelyesMiniMax-H3 only
contentyestext, image_url, video_url, audio_url items
durationyes4 to 15 seconds
rationoadaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
callback_urlnowebhook for status updates

Does Sume expose anything like it?

In the pages I read, no. Sume's /v1/videos takes prompt, duration, resolution, aspect_ratio, frame_images, input_references, generate_audio and callback_url. provider.options is rejected with 400 unsupported_parameter because the passthrough list is empty in v1 for every model, and seed and size are rejected as well. There is no endpoint that returns only a prompt.

That also means MiniMax's extra generation options, such as prompt expansion mode, are not reachable through Sume. If you need the enhancer specifically, you would call MiniMax directly; Sume will not proxy that call.

What is the Sume-side workflow instead?

Do the context reading yourself. Describe each reference by role in the prompt, send the references as input_references or Video Router reference_*_urls, and keep the prompt under MiniMax's 7,000 characters. Sume's guide says reference-to-video treats images as visual guidance, not exact frames, and that frame_images takes precedence if you send both. Audio and video references are honored by minimax-h3 and minimax-h3-max.

A concrete pattern: ask your own assistant to draft the prompt from your assets, review it, then submit once. You pay for one video, not for a draft video.

import os, requests

prompt = open("reviewed-prompt.txt").read().strip()
assert 0 < len(prompt) <= 7000
r = requests.post(
    "https://api.sume.com/v1/video-router/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "h3-ctx-001"},
    json={"model": "minimax-h3", "prompt": prompt, "duration": 8,
          "resolution": "768p",
          "reference_image_urls": ["https://example.com/product.png"],
          "mode": "async"},
)
print(r.status_code, r.json())

When is skipping the enhancer a loss?

If your inputs are a mix of long audio and video and your prompts come out vague, a vendor-side reader may catch details you skip. The honest trade is control and a single bill on Sume against an extra vendor hop elsewhere. For the writing side, see the H3 negative-direction prompt guide, and for what the switch does and does not do on Sume, the prompt expansion post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume