Reel series: keep one product's look across five AI video episodes
How to keep one product looking the same across a five-episode Reel series by pinning a frame or passing references on Sume, and which models accept which.

To keep one product looking the same across five episodes, give every job the same product image: pin it as frame_images first_frame when the product must appear exactly as in the photo, or pass it as input_references when you want the model to follow its look but have freedom in the shot. Both fields are in POST /v1/videos, and reruns of the same prompt will still differ.
We did not read an Instagram page on a Reel Series feature, so this post is about the Sume side of the problem only. For a 9:16 Reel, Meta's Reels ads page (read 2026-10-07) recommends the vertical format.
Which field to use
The docs separate the two modes, and mixing them has a rule.
| Field | What the model does | Use it for |
|---|---|---|
| frame_images (first_frame / last_frame) | Starts or ends on the exact image; image-to-video | The product pack shot that must be exact |
| input_references | Treats the images as visual guidance, not exact frames; reference-to-video | Style and look across shots |
| Both in one request | frame_images controls the mode and the job runs as image-to-video | Avoid; pick one |
A series recipe
Keep what repeats in one place and change only the story beat.
- One product image, one style image and one fixed paragraph of look notes (colors, lighting, lens) that you paste into every prompt.
- A new beat per episode in one sentence: unboxing, use, close-up, result, call to action.
- The same model, resolution and aspect ratio for all five. Use
GET /v1/videos/modelsto confirm 9:16 is listed. - A stable
Idempotency-Keyper episode, so a retry cannot bill a second job.
Which models take references
From the video guide: the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references as well as images. Gemini Omni Flash 1.1 accepts image and video references but no audio. Check supported_input_references per model before you write the request, since a type the model does not list will be refused.
Check the look across episodes
After the five jobs finish, run video inspect with frames: { at: [0.5] } on each file, and put the five stills side by side. Check that the pack, label text and colors match. If one episode drifts, rerun only that job with a shorter prompt and the same product image. No model takes a seed, so you cannot reproduce a good take, and you should keep the file of any take you like.
Cost and limits
Each episode is its own job, and the poll response shows usage.cost, so add up five numbers rather than guessing. For per-tier episode costs of an avatar series, see the ten-episode mini-series test. For the field rules in detail, see frame_images and input_references together.
Sources
Related posts
More in Media tools
- Reels image ad with auto music, or your own video? Still to 9:16
Meta's Advantage+ Creative can add music to Reels image ads. When to instead turn the still into a 9:16 video with image-to-video on Sume.
- Reels picture-in-picture test: check the four corners first
Instagram is reported to be testing picture-in-picture. Crop each corner of your 1080x1920 Reel with video-filter, free to check, $0.02 to encode.
- Mute a clip or extract its audio? Trim audio drop vs audio detach
Video trim with audio drop returns a silent MP4; audio detach returns a separate wav or mp3 and leaves the video alone. Which to use, and the warnings.
- Repurposing TikTok and Reels to Shorts: restyle captions once, $0.20
Reposting a clip to YouTube Shorts? Keep a clean master and burn captions with Sume: $0.20 a job for up to 60 seconds, and restyles skip a second transcription.
Written by Sume