Virtual try-on API: which Sume call returns an image, which a video
Need a try-on photo or a try-on clip? On Sume the two catalog try-on Formats return video; a still comes from the image API. The table, plus one call for each.

It depends on what you want back. Sume's two try-on Formats, sume-virtual-try-on and sume-virtual-fitting, are declared as text-in, video-out, so they return a 9:16 clip. For a try-on still, call POST /v1/images with openai/gpt-image-2.5, a person photo and a garment photo as references. There is no single try-on endpoint that returns either on request.
The launch of ChatGPT Try On on October 1 2026 put the picture-first version in front of shoppers, so sellers now ask for both. Here is how to pick without guessing.
What do the two try-on Formats return?
Both appear on the Formats by Sume catalog as callable slugs under the reserved sume handle, and any key with formats:write may call one. In Sume's own registry, both declare an io of input_kind: text, output_kind: video; the docs say that io is the declared profile for picking a Format without calling it, and that it is null for Formats saved before registration. The registered example for sume-virtual-try-on builds a first frame with ChatGPT Image 2, then animates the outfit reveal with Seedance 2.5; sume-virtual-fitting builds the same way for a silhouette and drape preview.
The run is a Format run, not a one-shot call. It returns a 202 receipt, and you read artifacts[] and primary_output_url once the status is completed, per Runs and results.
When is the image API the better call?
When the deliverable is a picture: a product-page shot, a marketplace image, an ad still. The Image API takes up to 16 references on the ChatGPT Image 2.5 models, returns a Sume-hosted data[].url, and blocks up to 30 seconds before it falls back to a 202 job. It gives you one pose and one frame, with no motion and no run to poll in the common case.
| Route | Returns | Call shape | Use it for |
|---|---|---|---|
| sume-virtual-try-on | Video (declared output_kind video) | POST /v1/formats/sume/sume-virtual-try-on/runs, poll the run | A creator changing into the garment |
| sume-virtual-fitting | Video (declared output_kind video) | POST /v1/formats/sume/sume-virtual-fitting/runs, poll the run | Silhouette and drape on a listing |
| openai/gpt-image-2.5 | Image URL | POST /v1/images, usually 200 inside 30 s | A still on a product page |
| seedance-2.5 with a first frame | Video | POST /v1/videos, poll polling_url | Animating a still you already approved |
What does each call look like?
A Format run takes the garment as an attachment plus an instruction, and needs an Idempotency-Key on create, as Create a Format run says. The image call takes input_references. Below is the Format call in Python; swap the URL and the instruction for your own garment.
Keep the idempotency key tied to the SKU and a version you bump on purpose, not to the clock.
import os
import requests
key = os.environ["SUME_API_KEY"]
body = {
"instruction": "9:16 virtual try-on of the attached jacket: a creator changes into it, same room.",
"attachments": [
{"type": "input_image", "image_url": "https://cdn.example.com/jacket.jpg"}
],
"generation_spend_cap_usd": 20,
}
r = requests.post(
"https://api.sume.com/v1/formats/sume/sume-virtual-try-on/runs",
headers={"Authorization": f"Bearer {key}", "Idempotency-Key": "sku-77-tryon-v1"},
json=body,
timeout=60,
)
r.raise_for_status()
run = r.json()["data"]
print(run["id"], run["status"], run["status_url"])Can you chain the image call into the video call?
Yes, and it is a common way to control the look. Make the still first, look at it, and only then spend on motion. Because a Format run builds its own first frame, chaining is only worth it when you want to approve the frame yourself: generate with openai/gpt-image-2.5, check the garment against your packshot, then pass the accepted still to seedance-2.5 as a first_frame on POST /v1/videos. That route is documented in Video generation, which lists 4 to 30 seconds for seedance-2.5.
The cost is two calls and two bills, but each is small and each can be retried on its own with its own idempotency key. If you would rather have one receipt and one webhook, call the Format and let it do both steps.
Which one should a small store start with?
Start with the still. It is one call, it returns in the common case inside the 30-second wait, and you can judge it in seconds. If it holds up against your packshot, a clip is the next step for the products where motion adds something, such as a coat that moves or a skirt that swings. Starting with the Format makes sense when you want the whole preview as one deliverable and are happy to let the run build its own first frame and pace.
Whichever you pick, set a spend cap and an idempotency key on day one. They cost nothing and they turn a retry loop from a risk into a routine.
What does a third-party try-on API do differently?
Some vendors sell a dedicated try-on endpoint that takes a garment image and a person image and returns an image. We have covered that shape in the Kling try-on post and the FLUX try-on post. Sume's equivalent for stills is the reference edit, and its equivalent for clips is the Format. Neither returns a fit score or a size, and consumer features such as ChatGPT Try On are reported to carry no size or fit guarantee either, per Retail Technology Innovation Hub.
- Pick by deliverable: clip means a Format, still means the image call.
- A Format run is billed against its own spend cap; set
generation_spend_cap_usdon every call. - Both routes need public HTTPS image URLs.
- Pick one mode per video request: a first frame, or reference images; the OpenAPI says
frame_imageswin when both are sent. - Neither route certifies size or fit.
Sources
Related posts
More in Developers
- Test a webhook endpoint before go-live: a Sume CI gate (Python)
Use POST /v1/webhooks/test-deliveries to fire a signed webhook.test at your deployed URL and fail the deploy unless it answers 2xx. Python script included.
- Zod 4 discriminated union for Sume job and run webhooks (TypeScript)
Parse Sume job.* and format.run.terminal webhooks with one Zod 4 discriminatedUnion: typed branches, degraded runs, oversized receipts. Tested with Zod 4.
- Zod 4 toJSONSchema to Sume output_schema: nullable, not optional
z.toJSONSchema works for a Sume Format output_schema if you use nullable instead of optional. A tested table of what passes and what the validator rejects.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
Written by Sume