Draft vs final render cost: a 6-second clip on five Sume video models
The final render costs 2.5x to 10x the draft on Sume, depending on the model. A 6-second clip priced at the cheapest and the top tier, per model.

The question
Draft-then-final saves money only if the draft is cheap enough relative to the final. The ratio differs a lot between models, because each Sume video model has its own per-second rows. This page prices the same 6-second clip at each model's cheapest listed row and at its top row, using Sume's rate table (provider list times 1.25).
The table
Prices are per 6-second job. Seedance 2.5 is derived from its per-1,000-token rate, so its 720p row stands in for the draft.
| Model | Draft row | Draft price | Final row | Final price | Ratio |
|---|---|---|---|---|---|
| Wan 3.0 | 480p | $0.38 | 1080p | $1.50 | 4.0x |
| Gemini Omni Flash 1.1 | 360p | $0.225 | 1080p | $1.125 | 5.0x |
| Gemini Omni Flash 1.1 | 360p | $0.225 | 4K | $2.25 | 10.0x |
| minimax-h3 | 480p | $0.375 | 4K upscale | $1.20 | 3.2x |
| minimax-h3-max | 480p | $0.375 | 1080p | $1.20 | 3.2x |
| Seedance 2.5 | 720p | $3.47 | 1080p | $8.53 | 2.5x |
Reading the ratios
Omni Flash has the widest spread: a 360p draft at $0.225 against a 4K final at $2.25. Wan 3.0 is a clean 4x between 480p and 1080p. The two MiniMax models are the narrowest at about 3.2x, and Seedance 2.5 only offers 480p, 720p and 1080p in the Video Router docs, so its draft-to-final step is about 2.5x from 720p to 1080p.
The ratio is not the whole story. One draft only saves money if you reject enough of them before the final.
A draft does not carry over
No v1 video model on Sume accepts seed; each model reports seed: false and rejects the field (Video generation docs). The final is a new job, a new reservation and a new take. A draft tells you the prompt and framing work; it does not promise the same motion at a higher resolution.
Each job is reserved at submit and captured on success, so the draft and the final are two separate charges in the usage ledger (Generation admission).
Practical rule
- Draft on the model you will finish with when you can; switching models between draft and final changes the look.
- Use Omni 360p only if the final is Omni, and the 10x ratio is worth it for the amount of exploration you need.
- Cap the draft phase by count, not just by money: ten 480p Wan drafts cost the same as 2.5 finals at 1080p.
A budget for ten concepts
Take ten concepts, a 6-second clip each. A draft-and-final plan on Wan 3.0 costs 10 x $0.375 for the drafts, $3.75, plus a 1080p final for the three you keep: 3 x $1.50 = $4.50. That is $8.25 in total. Finals only, all ten at 1080p, would be $15.00. The draft stage pays for itself if it prevents at least three bad finals.
On Omni Flash the numbers are more dramatic: ten 360p drafts at $0.225 are $2.25 and three 4K finals at $2.25 are $6.75, $9.00 in total, against $22.50 for ten blind 4K renders.
The ratios in the table are the useful part, because they do not depend on how many clips you make. A model with a wide spread between its lowest and highest tier rewards drafting; a model whose rows are close together does not. Wan 3.0 at 480p against 1080p is a four to one spread, and Omni Flash from 360p to 4K is ten to one, so those are the two places where a cheap first pass pays off most.
Remember what a draft does not do. Without a seed, a final render is a new generation, so a draft tells you whether the prompt and framing work, not what the final frame will look like in detail. Use drafts to filter prompts, then budget finals as independent jobs and keep a margin for the occasional second take.
Sources
Related posts
More in Comparisons
- Eleven v4 tags like [light rain] vs a separate sound bed on Sume
Eleven v4 puts effects such as [light rain] inside the speech. Sume keeps voice and bed separate, with gain_db, loop and duck_db. The trade-off for editing.
- Directing delivery: Eleven v4 audio tags vs Gemini TTS style
Eleven v4 puts direction inline as audio tags like [laughs]; Gemini TTS adds a separate style field. Sume sends transcript, voice and language only.
- Eleven v4 IPA support vs Sume's pronunciation_dict_id for brand names
ElevenLabs says Eleven v4 improves IPA phoneme support. Sume TTS has an optional pronunciation_dict_id. How each helps with brand names today.
- Voice agent TTS latency: Eleven v4 Turbo 150ms vs Cartesia Sonic
Eleven v4 Turbo claims about 150ms to first speech; Cartesia Sonic-3.6 claims under 90ms. Not like-for-like; Sume TTS is not for live agents.
Written by Sume