Reel series: keep one product's look across five AI video episodes

How to keep one product looking the same across a five-episode Reel series by pinning a frame or passing references on Sume, and which models accept which.

5 min readSume
All posts

To keep one product looking the same across five episodes, give every job the same product image: pin it as frame_images first_frame when the product must appear exactly as in the photo, or pass it as input_references when you want the model to follow its look but have freedom in the shot. Both fields are in POST /v1/videos, and reruns of the same prompt will still differ.

We did not read an Instagram page on a Reel Series feature, so this post is about the Sume side of the problem only. For a 9:16 Reel, Meta's Reels ads page (read 2026-10-07) recommends the vertical format.

Which field to use

The docs separate the two modes, and mixing them has a rule.

frame_images versus input_references (read 2026-10-07)
FieldWhat the model doesUse it for
frame_images (first_frame / last_frame)Starts or ends on the exact image; image-to-videoThe product pack shot that must be exact
input_referencesTreats the images as visual guidance, not exact frames; reference-to-videoStyle and look across shots
Both in one requestframe_images controls the mode and the job runs as image-to-videoAvoid; pick one

A series recipe

Keep what repeats in one place and change only the story beat.

  • One product image, one style image and one fixed paragraph of look notes (colors, lighting, lens) that you paste into every prompt.
  • A new beat per episode in one sentence: unboxing, use, close-up, result, call to action.
  • The same model, resolution and aspect ratio for all five. Use GET /v1/videos/models to confirm 9:16 is listed.
  • A stable Idempotency-Key per episode, so a retry cannot bill a second job.

Which models take references

From the video guide: the Seedance 2.x models, Wan 3.0, MiniMax H3 and MiniMax H3 Max accept audio and video references as well as images. Gemini Omni Flash 1.1 accepts image and video references but no audio. Check supported_input_references per model before you write the request, since a type the model does not list will be refused.

Check the look across episodes

After the five jobs finish, run video inspect with frames: { at: [0.5] } on each file, and put the five stills side by side. Check that the pack, label text and colors match. If one episode drifts, rerun only that job with a shorter prompt and the same product image. No model takes a seed, so you cannot reproduce a good take, and you should keep the file of any take you like.

Cost and limits

Each episode is its own job, and the poll response shows usage.cost, so add up five numbers rather than guessing. For per-tier episode costs of an avatar series, see the ten-episode mini-series test. For the field rules in detail, see frame_images and input_references together.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume