Omni Flash 1.1 vs Wan 3.0: a 10-second clip at 1080p, $1.875 vs $2.50
On Sume a 10-second clip is $1.25 at 720p on both gemini-omni-flash-1.1 and wan-3.0, and $1.875 vs $2.50 at 1080p. Omni caps at 10 s and always adds audio.

A 10-second 720p clip costs $1.25 on both gemini-omni-flash-1.1 and wan-3.0 on Sume, since both are $0.125 a second. At 1080p, Omni is cheaper: $1.875 against $2.50. Ten seconds is exactly Omni's maximum, so this is the longest clip you can make on that row, while Wan 3.0 runs to 30 seconds.
The grid
Omni lists 360p, 720p, 1080p and 4K, while Wan 3.0 lists 480p, 720p and 1080p. The rows below pair the tiers that exist on both and then show the extras.
| Tier | Omni Flash 1.1 per second | Omni 10 s | Wan 3.0 per second | Wan 3.0 10 s |
|---|---|---|---|---|
| Lowest tier (360p / 480p) | $0.0375 | $0.375 | $0.0625 | $0.625 |
| 720p | $0.125 | $1.25 | $0.125 | $1.25 |
| 1080p | $0.1875 | $1.875 | $0.25 | $2.50 |
| 4K | $0.375 | $3.75 | not listed | not listed |
What differs besides price
Per Sume's docs, Omni is 16:9 or 9:16 only, accepts up to 10 reference images and up to 3 reference videos of at most 3 seconds each, can edit a source video, and always generates native audio (generate_audio false is rejected). Wan 3.0 accepts 2 to 30 seconds and takes audio and video references.
- If you need a clip longer than 10 seconds, Omni is out.
- If you need silent output, Omni is out. Native audio cannot be turned off.
- If you need 4K, Wan 3.0 is out; Omni's 4K is $0.375 a second.
Same request, two models
Only the model, duration and resolution change.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: omni-10s-1080p" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "A vertical product clip on a desk, natural light",
"resolution": "1080p",
"duration": 10,
"aspect_ratio": "9:16",
"mode": "async"
}'Budget view
A hundred 10-second 1080p clips are $187.50 on Omni and $250.00 on Wan 3.0, so Omni saves $62.50 per hundred at 1080p. At 720p the two tie at $125.00. The savings case for Omni therefore rests on the 1080p tier, and the case for Wan 3.0 rests on length, silence or audio control.
The 360p draft trick
Omni's lowest tier is 360p at $0.0375 a second, so a 10-second draft is $0.375, which is 40 percent cheaper than Wan 3.0 at 480p ($0.625). If you want to test a prompt before paying for 1080p, 10 drafts on Omni 360p cost $3.75, and 10 finals at 1080p cost $18.75. The whole loop is $22.50, against $25.00 for 10 Wan 3.0 clips at 1080p with no cheaper draft tier than 480p.
This is a price comparison only. Whether a 360p Omni draft predicts the 1080p result is something to test on your own prompts.
Sources
Related posts
More in Comparisons
- Laugh and sigh tags in TTS: Gemini 3.8 tags vs Sume's emotion field
Gemini 3.8 TTS takes inline tags such as laugh, sigh and breath. Sume TTS documents an emotion hint, speed and volume instead; test tags on one short job.
- Generate at 1K and upscale, or generate 4K? Sume price check
Nano Banana 2.1 at 1K plus a Sume image upscale is $0.30, against $0.20 for native 4K. GPT Image 2.5 low plus upscale is $0.22 against $0.22 for high 4K.
- GPT Image 2.5 on Sume: low 1K $0.02475 vs high 4K $0.2225
GPT Image 2.5 spans 9x in price on Sume: $0.02475 for low at 1K, $0.055625 medium at 2K, $0.2225 high at 4K. Omit quality and the default is high.
- Griffin-Lite's 0.43 s latency: what a recorded avatar clip gives up
Tavus reports Griffin-Lite video latency of 0.43 s on average (0.27-0.59 s) in a research preview. A recorded clip on Sume is a job, so choose by use case.
Written by Sume