A 60-second ad with 4 b-roll clips and 8 stills: $14.99
One 60-second spokesperson ad on Sume: avatar standard tier, 4 Wan 3.0 b-roll clips, 8 GPT Image 2.5 stills and a new avatar come to $14.99.
One 60-second spokesperson ad costs $14.99 on Sume when it uses a new avatar on the standard tier, 4 five-second Wan 3.0 b-roll clips at 720p and 8 medium-quality GPT Image 2.5 stills. The avatar video is most of it.
Line items
Every line is a Sume list price read on 2026-10-09: Avatar Video standard at $0.184 per second, avatar creation at $0.95, Wan 3.0 720p five-second clips at $0.63 each, and medium 16:9 stills at $0.06 each.
| Item | Quantity | Unit price | Subtotal |
|---|---|---|---|
| New avatar | 1 | $0.95 | $0.95 |
| Avatar video (standard) | 60 s | $0.184/s | $11.04 |
| Wan 3.0 720p clips | 4 x 5 s | $0.63 | $2.52 |
| GPT Image 2.5 medium stills | 8 | $0.06 | $0.48 |
| Total | $14.99 |
Reusing the avatar
The avatar is a one-time cost. Ten ads in the same format reuse it, so ten cost $141.35 rather than $149.90. Narration is not a separate line because the avatar video speaks the script.
Each job reserves its price at submit and refunds it on failure, so a failed job does not change the total above. Avatar Video accepts scripts of 4 to 60 seconds.
Sources
Related posts
More in Use cases
- One 8:1 image sliced into 8 square carousel slides: $0.15 vs $1.20
Generate one 8:1 Nano Banana 2.1 image on Sume at $0.15 (2K), cut it into 8 squares for a swipe carousel. Cost vs 8 separate images, slicing code, seam gotchas.
- One 9-second dance on four brand avatars: Kling motion control, $5.67
Four avatars, one 9-second driving video: four Kling 3.0 Motion Control jobs at $0.1575 a second, 4 x 9 x 0.1575 = $5.67 on Sume. Request shape and queueing.
- One model for 1:1, 4:5, 9:16, 16:9 and 21:9: 6 Sume rows do it
Of 10 Sume image rows checked, only six list all of 1:1, 4:5, 9:16, 16:9 and 21:9. FLUX.2 pro does it at $0.0375, so five placements cost $0.1875.
- Photo avatar of a real person: when YouTube's AI label applies
A photo avatar scripted to say new words falls under YouTube's 'real person says something they did not' test. How Sume's three avatar inputs map to it.
Written by Sume