Outfit reveal clip: Seedance 2.5 or Gemini Omni Flash 1.1 on Sume?
Pick the Sume video model for a try-on or outfit reveal by the limits that matter: duration, resolution, references, audio and aspect ratio.

Use seedance-2.5 when you want a longer clip, up to 30 seconds, or a choice of 480p, 720p or 1080p; use gemini-omni-flash-1.1 when you want a short 3 to 10 second reveal with a start and end frame, up to 4K, and native audio. Both run through Sume's Video Router. Neither is better in general; the limits decide.
The reason to ask now is ChatGPT's new Try on button for clothing and accessories (read 2026-10-03, OpenAI help page), which pushes shops to show items on a body. A still is the start; a short reveal clip is the next ask, and the model you pick sets your cost and your constraints.
The limits side by side
These come from the Video Router doc and the first-party Format description. The Sume sume-virtual-try-on Format itself builds its first frame with ChatGPT Image 2 and animates with Seedance 2.5 reference-to-video, so Seedance 2.5 is the model Sume's own try-on recipe uses.
| Property | seedance-2.5 | gemini-omni-flash-1.1 |
|---|---|---|
| Duration | 4 to 30 seconds | 3 to 10 seconds |
| Resolution | 480p, 720p, 1080p | 360p, 720p, 1080p, 4K |
| Aspect ratio | Set with aspect_ratio, such as 9:16 | 16:9 or 9:16 |
| Image input | Documented reference-to-video use in the try-on Format | image_url plus optional end_image_url; up to 10 reference images |
| Audio | Check the catalog for the model | Always on; generate_audio: false is rejected |
| Edit an existing clip | Not described for this model in the docs | video_url edit mode |
| Billing | Provider list x 1.25 per output second | Provider list x 1.25 per output second |
How to choose
If the reveal is one move, a turn or a step, and you have both a start and an end still, gemini-omni-flash-1.1 with image_url and end_image_url is a direct fit, and the start and end frame post walks through it. If you want a 15 or 20 second piece with several beats, or you need 1080p without going to 4K prices, Seedance 2.5 is the model with the length.
Audio is the quiet difference. Gemini Omni Flash 1.1 always produces audio and does not let you turn it off. If your destination autoplays on mute that does not matter, but if your brand needs a silent file, plan to strip it downstream or pick another model. The Video Router catalog at GET /v1/video-router/models lists what each model accepts, so read it for the exact fields before pinning a model.
curl -s https://api.sume.com/v1/video-router/models \
-H "Authorization: Bearer $SUME_API_KEY"
# Same reveal, two models. Change only the model and the limits it allows.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: reveal-1042-seedance-v1" \
-d '{"model":"seedance-2.5","prompt":"The model turns and the jacket settles. Soft daylight.","resolution":"720p","duration":8,"aspect_ratio":"9:16","mode":"async"}'A fair test
Pick one try-on still and run both models with the same wording, at 720p, for the shortest duration both allow. Look at five things: does the garment keep its shape, does the print survive the move, are hands clean, does the face stay the same, and does the motion look like what you asked for. Read usage.cost on each and record the price per second you pay.
Do the test with your own product, not a demo, because garment type changes the result: knitwear, shiny fabric and long skirts each fail differently. After one round you will know which model to use for which category, and the decision is yours to make from your own samples rather than from any claim in this post. For the wider map of which Sume call gives which output, see the which-call-returns-which post, and for seeding a clip from a try-on still see the seed-frame post.
What this post cannot tell you
The Sume docs list limits and capabilities, not image quality, and this post has no benchmark. It does not say one model keeps garments truer than the other, because that has to be measured on your own products. If you need a defensible choice, make a small test set of ten garments across at least three fabric types, run both models, and have two people score the results blind. Then write the decision and its date next to the model id, since catalogues change and a choice made today may not hold in six months.
Sources
Related posts
More in Comparisons
- Plainly 50 minutes for $48 vs Sume Timeline at $0.10 a minute
Plainly Starter lists 50 export minutes for $48 a month billed yearly, about $0.96 a minute. Sume renders an output minute for $0.10. Reel math for 3 minutes.
- Pocket TTS voice cloning: a wav in, and what Sume does instead
Pocket TTS clones from a wav file you pass to --voice, with consent rules in its model card. Sume's API takes voice ids, not audio. Here is the difference.
- Quso.ai credits at about one per minute vs Sume priced by step
Quso.ai says one credit is roughly one minute processed, from 100 credits for $29. Sume prices steps: $0.01 transcript, $0.02 trim, $0.10 render, $0.20 caption.
- Qwen Image vs Qwen Image Max on the Sume image API
Qwen Image lists $0.025 and edits with up to 10 references; Qwen Image Max lists $0.09375 and is text-to-image only. Catalog table and request examples.
Written by Sume