Seedance 2, Omni or Wan 3.0 for 9:16 ads on Sume: how to choose
Three Sume video models side by side for vertical ads after Sora: clip length, resolutions, 9:16 in the docs, references, and the price of a 5 s clip.

For a 9:16 ad of 10 seconds or less, start with Gemini Omni Flash 1.1; for 15 seconds, Seedance 2.0; for 30 seconds in one request, Wan 3.0 or Seedance 2.5. The first two list 9:16 in the Sume docs; for Wan and Seedance 2.5 read supported_aspect_ratios first. A tracker records that the OpenAI Sora API ended on 2026-09-24 (Magic Hour tracker, read 2026-10-06), so this is a choice among models Sume lists, with the facts the Sume docs give.
No model is best at everything; this table is a way to filter, and your own twenty prompts are the test.
Side by side
| Question | Gemini Omni Flash 1.1 | Seedance 2.0 | Wan 3.0 |
|---|---|---|---|
| Duration | 3-10 s | 4-15 s | 2-30 s |
| Resolutions | 360p, 720p, 1080p, 4K | 480p, 720p, 1080p | 480p, 720p, 1080p |
| 9:16 stated in the docs | Yes (16:9 and 9:16 only) | Yes (six ratios) | Read supported_aspect_ratios |
| Reference types | Image and video, no audio | Image, video and audio | Image, video and audio |
| Native audio | Yes, synced, no toggle | Yes, generate_audio | Yes |
| 5 s clip, 720p | $0.63 | Priced per video token | $0.63 |
How to choose
- Short hook, vertical feed, sound on: Omni. It is the only one of the three whose whole ratio set is 16:9 and 9:16.
- A subject that must match a reference clip or audio: Seedance 2.0 or Wan 3.0, which accept video and audio references. Omni takes image and video references only.
- A single 20 to 30 second piece: Wan 3.0 (2-30 s). Seedance 2.5 (4-30 s) is the other option.
- A first and last frame: check
supported_frame_imageson the model you pin; Omni supports image-to-video with an end frame in the router doc.
Run all three on one brief
The loop sends the same prompt to each model at 720p and 9:16 and prints the job ids. Check supported_aspect_ratios first for Wan; a refused ratio comes back as a 400 and costs nothing.
import os
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
brief = "A runner ties her shoes at sunrise on a rooftop, handheld, warm light"
for model in ("gemini-omni-flash-1.1", "seedance-2", "wan-3.0"):
r = requests.post("https://api.sume.com/v1/videos", headers=H, timeout=30, json={
"model": model, "prompt": brief, "aspect_ratio": "9:16",
"resolution": "720p", "duration": 5})
print(model, r.status_code, r.json().get("id") or r.json())Judge the three clips on the first second, the hands and faces, and the sound. Then check cost: Omni and Wan are $0.63 for 5 seconds at 720p; Seedance is priced per video token, so read its quote from the catalog. The replacement-by-need post and the eight-model comparison give a wider view, and Seedance vs Sora 2 covers one pairing in detail.
Mistakes that skew a three-model test
Comparisons of generation models go wrong in boring ways. The most common is changing two things at once, such as the model and the duration, then crediting the model for the difference. Keep every field fixed except the model id: same prompt, same ratio, same resolution, same duration, and the same number of attempts per model.
The second is judging on one lucky or unlucky sample. Run each model on several prompts from your real brief, not one. The third is skipping the sound: Omni and Wan generate audio natively and Seedance has a generate_audio option, so a silent review can miss the very thing that separates them.
Keep the result of the comparison with its date and the prompts you used. A choice made in October on twenty prompts is a fact about October and those prompts, and the next time a model changes you will want the old numbers to compare against rather than a memory of which one felt better.
- Judge the first second of each clip first; it decides whether a viewer stays.
- Look at hands, text on screen and fast motion, which are where models differ most.
- Record cost per keeper, not cost per clip, as in the budget post.
- Pick one default and one fallback, and put both in config with the limits you checked.
Sources
Related posts
More in Comparisons
- Sync.so $0.05 a second vs Sume H3 Max lip sync: rates side by side
Sync.so lists $0.05 to $0.04 a second on four plans; Sume H3 Max lip sync derives $0.0625, $0.10 and $0.20. What each rate buys.
- Synthesia Starter's 12 minutes vs twelve 60-second Sume avatar jobs
Synthesia lists $29 for about 12 minutes and $89 for about 60. How those minutes map to Sume avatar jobs capped at 60 seconds, and what each setup is for.
- Tavus Starter vs Growth: break-even is about 1,014 minutes a month
Tavus Starter is $59 plus $0.37 a minute over 100; Growth is $397 with 1,250 minutes. The crossover, and when a rendered Sume clip fits better than minutes.
- Cost per minute of AI narration: MAI-Voice vs Sume at 900 characters
A minute of narration is about 900 characters. That is $0.0198 on MAI-Voice-2.1, $0.0135 on Flash and $0.0428 on Sume TTS, before the 5.5% fee. Worked table.
Written by Sume