Yoga studio class promo on Wan 3.0: 38 cents to $1.50 for 6 seconds
A 6-second yoga class promo on Sume's wan-3.0 costs 38 cents at 480p, 75 cents at 720p and $1.50 at 1080p. Pick the tier by where the clip will play.

A 6-second class promo on wan-3.0 costs $0.38 at 480p, $0.75 at 720p and $1.50 at 1080p on Sume. For a studio that posts a story or reel a few times a week, 720p at 75 cents is the usual stopping point, and 480p is enough to test a prompt before you pay for more.
Sume bills the provider list price times 1.25 and rounds the total up to the cent. Wan 3.0 lists at $0.05, $0.10 and $0.20 per second for 480p, 720p and 1080p, so Sume's rates are $0.0625, $0.125 and $0.25 per second (video router docs).
What a 6-second promo costs
The math is duration times the per-second rate, rounded up. Six seconds is a comfortable length for one pose flow or one studio shot with a title card added later.
| Resolution | Sume rate per second | 6-second clip | Four clips a week for a month (16) |
|---|---|---|---|
| 480p | $0.0625 | $0.38 | $6.08 |
| 720p | $0.125 | $0.75 | $12.00 |
| 1080p | $0.25 | $1.50 | $24.00 |
Start from a photo of your room
Text-only prompts for a yoga class tend to invent a room that is not yours. wan-3.0 supports image-to-video with a first frame, so send a photo of your actual studio as frame_images with frame_type: first_frame, and describe only the motion: a teacher folding forward, morning light moving across the mat.
If you want the same teacher or props across several clips, send input_references instead. The model uses reference images as visual guidance, not as exact frames. If you send both fields, frame_images decides the mode and the request runs as image-to-video (video generation docs).
- Photo of the room as first frame: keeps the space recognizable.
- Reference images: keeps a person or prop consistent, but not pixel-exact.
- Audio is on by default for this model; set
generate_audioto false if you will add your own music.
Draft at 480p, finish at the tier you need
The cheap way to work is to run a 480p draft at $0.38, read the motion, and only then re-run the final at 720p. Two drafts and one final at 720p is $1.51 for one promo.
Pick 1080p only if the clip will be shown on a studio screen or embedded on a landing page. For phone-first social posts the extra 75 cents per clip buys little.
Limits worth knowing
The model accepts 2 to 30 seconds, so a 6-second ask is well inside the range. Check GET /v1/videos/models for the aspect ratios the model lists before you ask for 9:16, and read the exact amount a job billed in usage.cost on the poll response.
Sume does not guarantee that a pose is anatomically correct. Review each clip before you publish it, especially hands and feet.
Practical notes
Why Wan 3.0 for this job: it takes a first-frame photo, runs at 480p, 720p and 1080p, and prices per second, which keeps a studio promo simple to budget. If you later want a longer class recap, the same id goes to 30 seconds, so one workflow covers a story, a reel and a short recap.
A practical routine: shoot one still per class type, keep the prompt library in a spreadsheet, and log usage.cost per job. After a month you will know your real cost per published clip, including rejected takes, which is the number that matters for budgeting.
Sources
Related posts
More in Use cases
- YouTube bumper ad 5-6 seconds: generate at 6 or trim on Sume
YouTube's help page puts bumper ads at 5 to 6 seconds. Which Sume video models can generate a 6 second clip directly, and when a $0.02 trim is the better route.
- YouTube Shorts ad under 60 seconds: stitch clips with Timeline
YouTube suggests Shorts ads under 60 seconds. Sume Timeline 1.0 stitches clips into one MP4 at $0.10 per output minute, so 60 seconds costs $0.10.
- AI album cover generator: square art at 3000×3000
Generate square album art, then upscale: Apple recommends at least 3000×3000. On Sume, generate 2400×2400 and upscale it 1.25× to reach 3000×3000.
- AI avatar for online course videos: build and update lessons
Use an AI avatar as your online course instructor: one reusable avatar, a short talking video per section, captions, and one Timeline join per lesson.
Written by Sume