Keep the same face from a GPT Image 2.5 still into an AI video

Use an approved GPT Image 2.5 still as a Sume video's first frame, then check later frames for face drift. What frame_images does and what it cannot promise.

5 min readSume
All posts

To carry one face from a GPT Image 2.5 still into a video on Sume, send the approved still as frame_images with frame_type: "first_frame" on POST /v1/videos. That fixes the opening frame, not every later frame, so check the rest of the clip for drift. For a talking presenter, create a reusable Avatar 1.0 from the photo instead.

Step one: lock the face in a still

Do the identity work in the image, where it is cheapest to retry. OpenAI's image prompting page says to list what must not change in a person edit, including face, features, skin tone and pose, and to restate the constraints if the result drifts. Use openai/gpt-image-2.5 on Sume for that, then host the approved image at a public HTTPS URL.

Step two: use it as the first frame

Sume's Video generation docs describe two ways to give a video model an image. frame_images takes a first or last frame for image-to-video. input_references gives style or content guidance for reference-to-video, which the model uses as visual guidance rather than exact frames. If you send both, frame_images takes precedence and the request is treated as image-to-video.

Only models that list first_frame in supported_frame_images accept it, so read GET /v1/videos/models before choosing a model.

Image inputs on Sume's video API (Sume docs; OpenAI Sora removal read 2026-10-02)
FieldModeWhat the image does
frame_images, first_frameImage to videoSets the opening frame.
frame_images, last_frameImage to videoSets the closing frame, where the model supports it.
input_referencesReference to videoVisual guidance, not an exact frame.

The request

Replace the model id with one that your catalog read shows accepting a first frame. The docs' example row is seedance-2, but read the live list rather than trusting an example.

curl -X POST "https://api.sume.com/v1/videos" \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "seedance-2",
    "prompt": "The woman in the first frame turns toward the window and smiles. Same face, same hairstyle, no cuts.",
    "frame_images": [
      {"type": "image_url",
       "image_url": {"url": "https://example.com/approved-still.png"},
       "frame_type": "first_frame"}
    ],
    "resolution": "720p"
  }'

What a first frame does not guarantee

A first frame anchors one image. Over a few seconds of head turns or lighting changes the face can still shift, and nothing in Sume's docs promises it will not. Pull frames from the finished clip and compare them to the still; the contact-sheet post shows how. Shorter clips and gentler motion give the model less room to drift.

Sora is no longer an option for this on OpenAI's side: its deprecations page lists the Sora 2 ids as removed on September 24, 2026.

If the video is a person speaking to camera, Sume's Avatar docs describe creating a reusable avatar from a photo input with input.type: "photo" and a public HTTPS image_url, then reusing its handle in avatar videos. Use a face you have the right to use.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume