Turn a ChatGPT try-on image into a video with Seedance 2.5
Saved a try-on image to your ChatGPT Library? Host it, then use it as the first frame of a 9:16 clip on seedance-2.5 through POST /v1/videos. A working script.

Download the try-on image, put it at a public HTTPS URL, and send it as the first_frame in frame_images to POST /v1/videos with model: "seedance-2.5". Sume does not read your ChatGPT Library, so the image has to leave ChatGPT first. The clip then starts from your picture and animates from there, for 4 to 30 seconds, per Sume's docs.
Use a photo of yourself or a person who agreed to it. A try-on image is a picture of a real face, and animating it is a bigger step than viewing it.
Where does the image come from?
ChatGPT Try On was announced on October 1 2026 and saves the generated image to a Library, per TechCrunch and the Neuron digest, both read on 2026-10-03. Neither describes an export API. The practical route is to save the picture from the app and upload it somewhere public: your own storage, a CDN, or a Sume asset. Check whether it carries a provenance mark before you publish; the watermark and C2PA post covers what to look for.
What does Seedance 2.5 take from Sume?
From the Video generation docs: seedance-2.5 accepts 4 to 30 seconds at 480p, 720p and 1080p. A request can send frame_images for first and last frames (image-to-video) or input_references for reference-to-video; the OpenAPI says frame_images take precedence when both are sent. Each entry needs a frame_type of first_frame or last_frame, and the image must be a public HTTPS URL.
| Setting | Value to send | Why |
|---|---|---|
| model | seedance-2.5 | Catalog id listed in the docs |
| frame_images | one entry, frame_type first_frame | The try-on image becomes frame one |
| aspect_ratio | 9:16 (match the image) | Check the model's supported_aspect_ratios |
| duration | 6 to 10 | Any whole second from 4 to 30 is accepted |
| resolution | 720p to start | Check the look before choosing 1080p |
| prompt | one small action | A turn or a step; long scripts drift |
How do you submit it and fetch the file?
Submit, poll the polling_url, then read unsigned_urls[0]. Video is asynchronous, so a script that exits after the POST has nothing. The docs example polls every 30 seconds; this one does the same and stops on failed.
Keep the first frame's aspect ratio equal to the output ratio you ask for, otherwise you are asking the model to invent the difference.
import os
import time
import requests
key = os.environ["SUME_API_KEY"]
h = {"Authorization": f"Bearer {key}"}
frame = {
"type": "image_url",
"image_url": {"url": "https://cdn.example.com/tryon.png"},
"frame_type": "first_frame",
}
body = {
"model": "seedance-2.5",
"prompt": "The person turns slowly to show the shirt from the side. Same room, same light.",
"frame_images": [frame],
"aspect_ratio": "9:16",
"resolution": "720p",
"duration": 8,
}
job = requests.post("https://api.sume.com/v1/videos", headers={**h, "Idempotency-Key": "tryon-clip-v1"}, json=body, timeout=60).json()
while True:
time.sleep(30)
s = requests.get(job["polling_url"], headers=h, timeout=60).json()
if s["status"] == "completed":
print(s["unsigned_urls"][0])
break
if s["status"] in ("failed", "cancelled"):
raise SystemExit(s.get("error", s["status"]))What can go wrong?
The first frame can be right and the garment can still drift by the end. Pull a few frames and compare them to your packshot; Sume's Video frames endpoint does that for a clip hosted on media.sume.com. The older first-frame check shows the t=0 case.
- Pick one mode per request: the first frame, or reference images. The OpenAPI says
frame_imageswin when both are sent. - A private or signed URL is rejected; the image must be reachable without auth.
- A try-on image of someone else needs their consent before you animate it.
- Label the clip as AI-generated where a platform asks for it.
Should you animate a try-on image at all?
Sometimes. A still tells a shopper how a shirt looks; a four-second turn tells them how the sleeve falls. If your goal is a product-page video, the cleaner path is to start from your own garment photo and a person you cast, rather than from an image a shopper-facing assistant made for one customer. A try-on in ChatGPT is reported to be built from that user's own reference photo, which makes it their likeness, not a stock asset.
If it is your own image, keep the prompt to one motion. Seedance 2.5 will happily invent more, and every invention is a chance for the garment to change. Eight seconds at 720p is enough to check whether the idea holds before you spend on a longer, higher-resolution take.
The same call works for any first frame, so the garment still, a generated model shot or a packshot on a person all go through one request. Sume's two apparel Formats, sume-virtual-try-on and sume-virtual-fitting, do the frame-then-animate steps for you in one run when you would rather not stitch the calls yourself.
Sources
Related posts
More in Media tools
- Turn a hum into music with AI: what Sume takes as input
Stability says hum-to-steer is coming. Sume's music API takes text and one optional image, not audio. Here is how to describe a hummed tune in a prompt.
- Vertical video subtitles: BBC's 3-line rule and Sume settings
BBC guidance for 9:16 subtitles: up to 3 lines, 90% width, placed a little high. How to set safe_width_ratio and anchor_ratio on Sume burned-in captions.
- Vimeo "Invalid Caption File": build clean WebVTT from Sume
Vimeo rejects a cue that starts at the previous cue's end. Sume sentence segments touch, so trim 1 ms, write UTF-8 WebVTT and upload. Script included.
- WebVTT cue settings (line, size) vs Sume anchor_ratio and width
WebVTT line:78%,center and size:90% have close matches in Sume's caption design fields. A converter script, and what per-cue settings a burned render loses.
Written by Sume