Midjourney V8.2 edit stills, animated on Sume
Finish a still in Midjourney V8.2, host it at a public HTTPS URL, and use it as the first frame of a Sume video request. Midjourney itself has no video id here.

Edit your still in Midjourney, put the final file at a public HTTPS URL, and send it to Sume as the first frame of a video request. On POST /v1/videos that is a frame_images entry with frame_type: first_frame; on the legacy Video 1.0 URL it is image_url.
What the Midjourney notes say
The Midjourney notes on Releasebot list a Sep 3, 2026 entry for a V8.2 edit model in the lightbox, with plain-instruction edits using up to 4 reference images. An Aug 27 entry describes the first V8.2 edit model with inpaint and outpaint, and a Sep 24 entry adds Korean language on midjourney.com. The Sume docs do not list a Midjourney model id, and this post does not cover generating inside Midjourney.
From still to first frame
Image inputs on Sume must be fetchable public HTTPS URLs. Localhost, private-network, non-HTTPS and signed or private URLs are rejected, as are mismatched content types. So download the Midjourney result and host it where Sume can fetch it.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mj-still-anim-001" \
-d '{
"model": "sume/auto",
"prompt": "Slow push-in; keep the subject sharp and the background drifting",
"aspect_ratio": "9:16",
"duration": 5,
"frame_images": [{
"type": "image_url",
"image_url": {"url": "https://example.com/final-still.png"},
"frame_type": "first_frame"
}]
}'Edit with references, then keep them
If you edited with several reference images in Midjourney, you can reuse that idea on Sume. input_references gives reference-to-video guidance rather than exact frames, and only models whose supported_input_references lists a type accept it. If both frame_images and input_references are sent, frame_images takes precedence.
| Step | Tool | Field or output |
|---|---|---|
| Edit the still | Midjourney V8.2 edit model | Up to 4 reference images per the notes |
| Host the file | Your storage | Public HTTPS URL |
| Animate | Sume POST /v1/videos | frame_images with first_frame |
| Read the clip | Sume jobs | GET /v1/jobs/{id}/result |
Poll the polling URL the submit returns, or GET /v1/jobs/{id}/status. Auto create controls default to 720p and 8 seconds, with 3 to 10 second clips at 16:9 or 9:16. Check GET /v1/videos/models for the exact limits of a model you pin.
Sources
Related posts
More in Use cases
- H3 mixing camera, character and audio from three inputs
MiniMax H3 can mix camera movement from one video, a character from an image and audio from another. Sume's minimax-h3 takes image, video and audio references.
- Mixed-language explainer: two TTS jobs joined into one track
A voice that handles 23 or 90 languages still needs one language per request. Render each language as its own job, then join the takes with timeline audio.
- Naver Clip video: a 9:16 Sume clip with Korean captions
Making a vertical 9:16 clip for Naver Clip? Generate it on Sume, then burn Korean captions with the korean-ad style and language ko.
- Netflix Ads: 10-75 s at 16:9 1920x1080, render 16:9 native
Netflix Ads video specs: 16:9 at 1920x1080, 10 to 75 seconds, H.264. Render 16:9 natively on Sume and sequence clips for longer spots.
Written by Sume