Microdrama AI video: keep the same characters across episodes
Sume's video API has no cast-lock switch. To keep a character across episodes, start every shot from the same still with first_frame and reuse references.

Sume's video API has no setting that locks a character's face across episodes, so consistency is something you build into the inputs: start each shot from the same still of the character with first_frame, pass the same images as input_references, and keep the prompt wording for that character identical from episode to episode. The API does not promise the result will match, so you review every shot.
That matters more now because YouTube says Shorts are being organized into "dedicated seasons and episodes" for microdramas and series, per its blog post, read 2026-09-29. A series is watched back to back, and viewers notice when the lead changes face. The Sume details are from the Video generation docs.
Which request fields carry a character from shot to shot?
Two image fields exist, and they do different jobs. Choose per shot, not per series.
If a request carries both fields, frame_images takes precedence and the request is treated as image-to-video, so do not expect references to add anything once you send a first frame.
| Field | Mode | Use it for a microdrama when |
|---|---|---|
frame_images with frame_type: first_frame | Image-to-video | The shot must open on an exact frame: the lead at the door, in the same costume. |
frame_images with frame_type: last_frame | Image-to-video | You want the shot to end on a known frame, such as the cliffhanger close-up. |
input_references | Reference-to-video | You want images as visual guidance for a character or set, not exact frames. |
Which models accept those fields?
Not every model does. Each entry from GET /v1/videos/models lists supported_frame_images and supported_input_references, and only a model that lists a type accepts it. Read those two arrays for the model you plan to use before you design the series around it. The docs give the shape of an entry: supported_frame_images can be first_frame and last_frame, and supported_input_references can be image_url, video_url and audio_url.
Vertical is a per-model check too. supported_aspect_ratios is where 9:16 shows up, and gemini-omni-flash-1.1 is listed at 3–10 seconds in 16:9 or 9:16. The Video Router docs say to read capabilities from GET /v1/video-router/models rather than assume one envelope.
What does one consistent shot request look like?
Use one hosted image of the character as the opening frame for every shot of an episode. The same URL, the same character sentence in the prompt, and a fixed aspect_ratio of 9:16.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ep02-shot03" \
-d '{
"model": "seedance-2",
"prompt": "Mira, short dark hair, green raincoat, opens the door and freezes",
"aspect_ratio": "9:16",
"frame_images": [{
"type": "image_url",
"image_url": {"url": "https://example.com/mira-front.png"},
"frame_type": "first_frame"
}]
}'Does a fixed seed help?
No. The Sume docs say no v1 video model accepts seed: each reports seed: false and rejects the field, so you cannot pin a look that way. Consistency has to come from the frame and reference images and from repeated prompt wording.
Sume does accept an Idempotency-Key on POST /v1/videos, and a replay returns the original job. That is for safe retries after a network error, not for getting a second take; see the Video generation docs.
How should I plan a series so the cast survives?
Fix the character images before episode one, name them in a short sheet, and reuse the same lines of prompt text for the character in every shot. Prefer image-to-video for shots where the face is large in frame and references for wide shots. Generate the whole episode's shots, look at them side by side, and rerun only the ones that drifted. For the beat structure of each episode, see the microdrama cliffhanger shot list, and for how Shorts series work on YouTube, the series rules.
Sources
Related posts
More in Use cases
- Nano Banana consistent characters and 14 objects: the Sume cap
Google says Nano Banana 2 keeps up to 5 characters and 14 objects consistent. Sume takes 10 references per request, so 19 subjects need a plan.
- Nano Banana character consistency: reference limits and a Sume request
Google documents character reference slots for its Nano Banana models. The limits, and how to keep a character across images with input_references on Sume.
- Fundraising video ideas for nonprofits, made with AI
Fundraising video ideas built on real material: one true story, a program update, the ask and the link. Where AI helps, where it must stop, and what it costs.
- On hold music for business: make the music and messages
On hold music for a business is a calm track, often with spoken messages, played to waiting callers. How to generate both, in phone-ready formats.
Written by Sume