MiniMax H3-Context-IR returns a prompt, not a video: Sume view
MiniMax's H3-Context-IR endpoint reads text, images, audio and video and returns an enhanced prompt without making a video. What Sume has instead.

H3-Context-IR is a MiniMax endpoint, POST /v2/h3_context_ir, that reads your text, images, video and audio and returns an enhanced video prompt without creating a video. Sume's docs list no equivalent standalone prompt-enhancer endpoint, so on Sume you write the prompt yourself or ask an agent to write it, then submit minimax-h3 with the finished text.
Everything about the MiniMax endpoint below is from its page, Create H3-Context-IR Task, read on 2026-10-02. Sume's side is from Video generation and Video Router.
What does H3-Context-IR actually do?
According to the page, it interprets multimodal context across text, images, audio and video and generates an enhanced video prompt, and it does not create a video generation task. The request takes model (only MiniMax-H3), a content array, a required duration of 4 to 15 seconds, an optional ratio, and an optional callback_url. It returns a task_id with task type h3_context_ir; you fetch the finished prompt from content.prompt through the query-task endpoint.
It uses the same media limits as video generation: images up to 30 MB, videos up to 50 MB and 3 clips, audio up to 15 MB and 3 clips, and 64 MB for the whole request. The text item is limited to 7,000 characters.
| Field | Required | Value on the page |
|---|---|---|
| model | yes | MiniMax-H3 only |
| content | yes | text, image_url, video_url, audio_url items |
| duration | yes | 4 to 15 seconds |
| ratio | no | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| callback_url | no | webhook for status updates |
Does Sume expose anything like it?
In the pages I read, no. Sume's /v1/videos takes prompt, duration, resolution, aspect_ratio, frame_images, input_references, generate_audio and callback_url. provider.options is rejected with 400 unsupported_parameter because the passthrough list is empty in v1 for every model, and seed and size are rejected as well. There is no endpoint that returns only a prompt.
That also means MiniMax's extra generation options, such as prompt expansion mode, are not reachable through Sume. If you need the enhancer specifically, you would call MiniMax directly; Sume will not proxy that call.
What is the Sume-side workflow instead?
Do the context reading yourself. Describe each reference by role in the prompt, send the references as input_references or Video Router reference_*_urls, and keep the prompt under MiniMax's 7,000 characters. Sume's guide says reference-to-video treats images as visual guidance, not exact frames, and that frame_images takes precedence if you send both. Audio and video references are honored by minimax-h3 and minimax-h3-max.
A concrete pattern: ask your own assistant to draft the prompt from your assets, review it, then submit once. You pay for one video, not for a draft video.
import os, requests
prompt = open("reviewed-prompt.txt").read().strip()
assert 0 < len(prompt) <= 7000
r = requests.post(
"https://api.sume.com/v1/video-router/generate",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "h3-ctx-001"},
json={"model": "minimax-h3", "prompt": prompt, "duration": 8,
"resolution": "768p",
"reference_image_urls": ["https://example.com/product.png"],
"mode": "async"},
)
print(r.status_code, r.json())When is skipping the enhancer a loss?
If your inputs are a mix of long audio and video and your prompts come out vague, a vendor-side reader may catch details you skip. The honest trade is control and a single bill on Sume against an extra vendor hop elsewhere. For the writing side, see the H3 negative-direction prompt guide, and for what the switch does and does not do on Sume, the prompt expansion post.
Sources
Related posts
More in Developers
- MiniMax task statuses vs Sume job statuses: a mapping table
MiniMax reports queued, running, succeeded, failed, cancelled. Sume job status uses queued, processing, completed, failed, canceled. Map them correctly.
- Mirage Tesseract needs local files; Sume needs public HTTPS URLs
Mirage Tesseract runs on local files in an agent environment. Sume avatar and lip-sync inputs must be public HTTPS URLs. How to hand clips between the two.
- Music API has no duration field: steer length in the prompt
Sume's music router rejects duration and duration_seconds. Ask for a 30-second or 2-minute track in the prompt, with section timestamps, and verify.
- Music API metadata field: tag a track job with your own ids
Sume's music request has an optional metadata field, stored on the job and not sent to the provider. Use it to match tracks to your own records.
Written by Sume