H3 mixing camera, character and audio from three inputs
MiniMax H3 can mix camera movement from one video, a character from an image and audio from another. Sume's minimax-h3 takes image, video and audio references.

MiniMax describes H3 as using unified context across text, images, video and audio, with an example that applies camera movement from one video while a character from an image sings with audio from another source. Sume's minimax-h3 accepts image, video and audio references, and audio must be paired with at least one image or video reference.
What MiniMax says
From the MiniMax H3 post, read on 2026-10-03: H3 takes a unified context of text, images, video and audio. The example given combines three inputs, each contributing something different.
| Input | Contributes |
|---|---|
| Reference video | Camera movement |
| Image | The character |
| Audio from another source | What the character sings |
What Sume exposes
Limits are per model. minimax-h3 accepts 5 to 15 seconds at native 480p and 768p; 768p is first-class, not 720p. Audio and video references are honored by the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max. Confirm what a model accepts by reading supported_input_references from GET /v1/videos/models.
In the flat Video Router fields, reference_audio_urls requires at least one image or video reference. So the audio-only case is not valid; the three-input mix is.
Submit the mix
This uses the Video Router wire with flat reference fields. All media must be public HTTPS URLs. Reuse the Idempotency-Key on retries.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-mix-001" \
-d '{
"model": "minimax-h3",
"prompt": "The character from the image sings, with the camera move of the reference video",
"reference_image_urls": ["https://example.com/character.png"],
"reference_video_urls": ["https://example.com/camera-move.mp4"],
"reference_audio_urls": ["https://example.com/song.mp3"],
"resolution": "768p",
"duration": 8,
"mode": "async"
}'Set expectations
A prompt describing all three roles helps, but the docs do not promise that a reference is copied exactly. Treat the output as a take to review. Read capabilities from the model catalog before assuming one envelope, since limits differ between H3, H3 Max and other models.
- Keep the audio within the clip length you request.
- Use a short test at 5 seconds before a longer run.
- Poll the job and read
usage.coston the receipt.
Sources
Related posts
More in Use cases
- Mixed-language explainer: two TTS jobs joined into one track
A voice that handles 23 or 90 languages still needs one language per request. Render each language as its own job, then join the takes with timeline audio.
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Netflix Ads: 10-75 s at 16:9 1920x1080, render 16:9 native
Netflix Ads video specs: 16:9 at 1920x1080, 10 to 75 seconds, H.264. Render 16:9 natively on Sume and sequence clips for longer spots.
- Nonprofit year-end impact recap from five photos for Giving Tuesday
A 50-second year-end recap for a nonprofit from five photos: voice-over, music bed, one Timeline render and captions. $0.4678 on Sume, from docs.
Written by Sume