Turn a ChatGPT try-on image into a video with Seedance 2.5

Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.

5 min readSume
All posts

Download the try-on image, put it at a public HTTPS URL, and send it as the first_frame in frame_images to POST /v1/videos with model: "seedance-2.5". Sume does not read your ChatGPT Library, so the image has to leave ChatGPT first. The clip then starts from your picture and animates from there, for 4 to 30 seconds, per Sume's docs.

Use a photo of yourself or a person who agreed to it. A try-on image is a picture of a real face, and animating it is a bigger step than viewing it.

Where does the image come from?

ChatGPT Try On was announced on October 1 2026 and saves the generated image to a Library, per TechCrunch and the Neuron digest, both read on 2026-10-03. Neither describes an export API. The practical route is to save the picture from the app and upload it somewhere public: your own storage, a CDN, or a Sume asset. Check whether it carries a provenance mark before you publish; the watermark and C2PA post covers what to look for.

What does Seedance 2.5 take from Sume?

From the Video generation docs: seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p. A request can send frame_images for first and last frames (image-to-video) or input_references for reference-to-video; the OpenAPI says frame_images take precedence when both are sent. Each entry needs a frame_type of first_frame or last_frame, and the image must be a public HTTPS URL.

Seedance 2.5 settings for a try-on clip, read 2026-10-03
SettingValue to sendWhy
modelseedance-2.5Catalog id listed in the docs
frame_imagesone entry, frame_type first_frameThe try-on image becomes frame one
aspect_ratio9:16 (match the image)Check the model's supported_aspect_ratios
duration6 to 10Any whole second from 4 to 30 is accepted
resolution720p to startCheck the look before choosing 1080p
promptone small actionA turn or a step; long scripts drift

How do you submit it and fetch the file?

Submit, poll the polling_url, then read unsigned_urls[0]. Video is asynchronous, so a script that exits after the POST has nothing. The docs example polls every 30 seconds; this one does the same and stops on failed.

Keep the first frame's aspect ratio equal to the output ratio you ask for, otherwise you are asking the model to invent the difference.

import os
import time
import requests

key = os.environ["SUME_API_KEY"]
h = {"Authorization": f"Bearer {key}"}
frame = {
    "type": "image_url",
    "image_url": {"url": "https://cdn.example.com/tryon.png"},
    "frame_type": "first_frame",
}
body = {
    "model": "seedance-2.5",
    "prompt": "The person turns slowly to show the shirt from the side. Same room, same light.",
    "frame_images": [frame],
    "aspect_ratio": "9:16",
    "resolution": "720p",
    "duration": 8,
}
job = requests.post("https://api.sume.com/v1/videos", headers={**h, "Idempotency-Key": "tryon-clip-v1"}, json=body, timeout=60).json()
while True:
    time.sleep(30)
    s = requests.get(job["polling_url"], headers=h, timeout=60).json()
    if s["status"] == "completed":
        print(s["unsigned_urls"][0])
        break
    if s["status"] in ("failed", "cancelled"):
        raise SystemExit(s.get("error", s["status"]))

What can go wrong?

The first frame can be right and the garment can still drift by the end. Pull a few frames and compare them to your packshot; Sume's Video frames endpoint does that for a clip hosted on media.sume.com. The older first-frame check shows the t=0 case.

  • Pick one mode per request: the first frame, or reference images. The OpenAPI says frame_images win when both are sent.
  • A private or signed URL is rejected; the image must be reachable without auth.
  • A try-on image of someone else needs their consent before you animate it.
  • Label the clip as AI-generated where a platform asks for it.

Should you animate a try-on image at all?

Sometimes. A still tells a shopper how a shirt looks; a four-second turn tells them how the sleeve falls. If your goal is a product-page video, the cleaner path is to start from your own garment photo and a person you cast, rather than from an image a shopper-facing assistant made for one customer. A try-on in ChatGPT is reported to be built from that user's own reference photo, which makes it their likeness, not a stock asset.

If it is your own image, keep the prompt to one motion. Seedance 2.5 will happily invent more, and every invention is a chance for the garment to change. Eight seconds at 720p is enough to check whether the idea holds before you spend on a longer, higher-resolution take.

The same call works for any first frame, so the garment still, a generated model shot or a packshot on a person all go through one request. Sume's two apparel Formats, sume-virtual-try-on and sume-virtual-fitting, do the frame-then-animate steps for you in one run when you would rather not stitch the calls yourself.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume