Brand character keyframe: GPT Image 2.5 refs, then first_frame
Build an opening frame that keeps your character on model with up to 16 reference images on GPT Image 2.5, then animate it by pinning it as the first frame.

For a brand character that looks the same in every video, make the opening frame as an image first. GPT Image 2.5 on Sume accepts up to 16 reference images, so a model sheet and a product shot can steer one keyframe. Then send that keyframe to POST /v1/videos as a first_frame in frame_images, so the clip begins on the approved image.
Why start with an image
A video model that starts from text invents the character each time. A video model that starts from an approved still keeps it. An image is also cheaper to redo than a clip, so you can reject five keyframes before you spend on one video.
Sume lists openai/gpt-image-2.5 (Flare) and openai/gpt-image-2.5-sunburst. Both support text-to-image, up to 16 image references, an optional mask_url and background: auto|transparent|opaque. The quality field takes auto, low, medium, high, xhigh or max, and the default is high.
Custom sizes and limits
image_size takes named presets, auto or custom pixels. For custom pixels, both edges must be multiples of 16 and the maximum edge is 3840. The ratio must be at most 3:1, and the image must hold between 655,360 and 8,294,400 pixels. A 1080x1920 vertical frame fits: both edges divide by 16 and the pixel count is 2,073,600.
| Rule | Value |
|---|---|
| Reference images | Up to 16 |
| Custom edge | Multiple of 16, max 3840 |
| Aspect ratio | At most 3:1 |
| Pixel count | 655,360 to 8,294,400 |
| Default quality | high |
Animate the approved frame
After you approve the keyframe and have a public HTTPS URL for it, pin it as the first frame. The catalog tells you which models accept first_frame, so check supported_frame_images for the model you choose.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: mascot-spot-001" \
-d '{
"model": "seedance-2",
"prompt": "The character waves, then holds up the product and smiles",
"duration": 6,
"aspect_ratio": "9:16",
"frame_images": [
{"type": "image_url",
"image_url": {"url": "https://example.com/mascot-keyframe.png"},
"frame_type": "first_frame"}
]
}'Keep the character stable
Use the same model sheet in every keyframe request, and keep the character description out of the video prompt. The picture already carries it. Spend the prompt on motion. When a new campaign needs a new outfit, change the references, not the whole workflow.
Sources
Related posts
More in Use cases
- Brand name misheard in YouTube auto captions: burn your own script
YouTube warns that auto captions can misread accents and noisy audio. For a brand name that must be right, send script_text to Sume and burn your own wording.
- Branded coach character repeats your phone-filmed move: $3.15 per 20 s
Film a move on your phone, give Kling motion control a still of your coach mascot, and the mascot repeats the move. 20 seconds bills $3.15 at $0.1575 a second.
- Budget a translated script per sentence before you dub it
Turn each source sentence's seconds into a character budget before translating, so the dubbed read fits the slot. A method with Sume sentence segments.
- Reel speed 0.5x to 3x: burn captions after the final speed change
Instagram lets creators change Reel speed from 0.5x to 3x. Burned-in captions speed up with the clip, so a 1.5 s line is 0.5 s at 3x. Burn them last.
Written by Sume