First and last frame to video with MiniMax H3: request and rules
MiniMax H3 fills the motion between an opening and a closing image. On Sume send frame_images with first_frame and last_frame; ratio, length and price rules.

To make MiniMax H3 animate between two images, send frame_images on POST /v1/videos with one entry typed first_frame and one typed last_frame; the model fills the motion between them, guided by your prompt. One first_frame alone gives image-to-video, none gives text-to-video, and Sume refuses a last_frame without a first_frame.
Modes are from MiniMax's model card and fal's guide; fields from the Sume Video generation docs, read 2026-09-29.
What does the request look like?
Each entry needs a frame_type. The prompt describes the motion, not the images.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: h3-frames-001" \
-d '{
"model": "minimax-h3",
"prompt": "The headline slides down into place and the car lights shift from dark to red.",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/start.png" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/end.png" }, "frame_type": "last_frame" }
],
"resolution": "768p",
"duration": 6
}'How many frame images can I send?
| Images sent | Result |
|---|---|
| 0 | Text-to-video |
| 1 | First-frame or last-frame video (Sume accepts the first-frame form only) |
| 2 | First-and-last-frame video |
Which aspect ratio does the clip use?
fal's guide says output follows the aspect ratio of the uploaded image for first and last frame. Sume's code sends no aspect ratio on the image-to-video route for these ids, so do not rely on aspect_ratio to crop; crop your images to the ratio you want.
What if I send frame_images and input_references together?
frame_images wins and the request is image-to-video, per the docs. Use input_references alone when you want reference-to-video, and note that a single reference image with no first or last frame is priced and routed as reference-to-video.
Sources
Related posts
More in Developers
- Multi-shot AI video prompts: MiniMax H3 shot labels and timecodes
MiniMax H3 models multiple shots natively. Write [Shot 1] labels or timecoded blocks in the prompt; syntax from the vendor guides and a Sume request.
- MiniMax H3 prompt guide: reference jobs, negative direction, length
The MiniMax H3 prompting rules that change output: give each reference a job, write timecodes, say what to avoid, and stay under 7,000 characters.
- MiniMax H3 API in Python: submit, poll and download a video job
A runnable Python example for MiniMax H3 on Sume: POST /v1/videos, poll the job until completed, then download the clip from unsigned_urls. Uses httpx.
- Video to video motion transfer with AI: MiniMax H3 reference video
MiniMax H3 lists V2V motion transfer. Send a motion video and character images as references on Sume, write each reference's job, and mind the limits.
Written by Sume