Keep the same face from a GPT Image 2.5 still into an AI video
Use an approved GPT Image 2.5 still as a Sume video's first frame, then check later frames for face drift. What frame_images does and what it cannot promise.

To carry one face from a GPT Image 2.5 still into a video on Sume, send the approved still as frame_images with frame_type: "first_frame" on POST /v1/videos. That fixes the opening frame, not every later frame, so check the rest of the clip for drift. For a talking presenter, create a reusable Avatar 1.0 from the photo instead.
Step one: lock the face in a still
Do the identity work in the image, where it is cheapest to retry. OpenAI's image prompting page says to list what must not change in a person edit, including face, features, skin tone and pose, and to restate the constraints if the result drifts. Use openai/gpt-image-2.5 on Sume for that, then host the approved image at a public HTTPS URL.
Step two: use it as the first frame
Sume's Video generation docs describe two ways to give a video model an image. frame_images takes a first or last frame for image-to-video. input_references gives style or content guidance for reference-to-video, which the model uses as visual guidance rather than exact frames. If you send both, frame_images takes precedence and the request is treated as image-to-video.
Only models that list first_frame in supported_frame_images accept it, so read GET /v1/videos/models before choosing a model.
| Field | Mode | What the image does |
|---|---|---|
| frame_images, first_frame | Image to video | Sets the opening frame. |
| frame_images, last_frame | Image to video | Sets the closing frame, where the model supports it. |
| input_references | Reference to video | Visual guidance, not an exact frame. |
The request
Replace the model id with one that your catalog read shows accepting a first frame. The docs' example row is seedance-2, but read the live list rather than trusting an example.
curl -X POST "https://api.sume.com/v1/videos" \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2",
"prompt": "The woman in the first frame turns toward the window and smiles. Same face, same hairstyle, no cuts.",
"frame_images": [
{"type": "image_url",
"image_url": {"url": "https://example.com/approved-still.png"},
"frame_type": "first_frame"}
],
"resolution": "720p"
}'What a first frame does not guarantee
A first frame anchors one image. Over a few seconds of head turns or lighting changes the face can still shift, and nothing in Sume's docs promises it will not. Pull frames from the finished clip and compare them to the still; the contact-sheet post shows how. Shorter clips and gentler motion give the model less room to drift.
Sora is no longer an option for this on OpenAI's side: its deprecations page lists the Sora 2 ids as removed on September 24, 2026.
If the video is a person speaking to camera, Sume's Avatar docs describe creating a reusable avatar from a photo input with input.type: "photo" and a public HTTPS image_url, then reusing its handle in avatar videos. Use a face you have the right to use.
Sources
Related posts
More in Use cases
- Korea AI label tiers: invisible watermark vs visible for deepfakes
Korea's AI Basic Act allows invisible watermarks for obvious AI content like animation, but deepfakes need a visible label. How to tell which tier a clip is.
- Korea AI Basic Act grace period: when do 30 million won fines start?
Korea's AI Basic Act took effect in January 2026 with a grace period of at least a year before fines of up to 30 million won. What to prepare now.
- Korea AI Basic Act prior notice: tell users the product uses AI
MSIT's guidelines name two duties: prior notice that a service uses generative AI, and labeling outputs. What prior notice looks like in an API product.
- Korea AI Basic Act: who must label AI video, provider or user?
Under Korea's AI Basic Act the duty rests on AI providers and operators, not private users. What that means for a business that renders video through an API.
Written by Sume