30-second UGC ad: Avatar 1.0 job or Seedance 2.5 reference job?
A talking-host 30-second ad is $7.35 on Avatar 1.0 at plus. A moving, reference-driven 30-second ad is a seedance-2.5 job (4-30 s). How to choose.
Should a 30-second UGC ad be an avatar job or a video-model job? If one person talks to the camera from a script, run Avatar 1.0: 30 seconds at the plus tier is $7.35. If the ad moves through places, cuts between products or follows a reference clip, run a seedance-2.5 job, which Sume lists at 4-30 seconds.
Thirty seconds is now a normal single-job length. ByteDance says Seedance 2.5 generates up to 30 seconds per pass with joint audio-video generation, so the old habit of stitching six 5-second clips is optional.
What each one does
The tradeoff is about control, not quality labels.
| Question | Avatar 1.0 | seedance-2.5 |
|---|---|---|
| Input | Script or video_inputs, avatar_handle | Prompt, frame images, references |
| Duration | 4-60 s per job | 4-30 s per job |
| Same face every time | Yes, by handle | Via reference images |
| 30 s cost | $5.52 / $7.35 / $16.50 (standard, plus, max) | Read pricing_skus from GET /v1/videos/models |
When the avatar wins
A testimonial, a founder explainer or a product pitch needs the same face and exact words. The avatar route takes the script literally, supports a silence beat inside video_inputs, and can burn captions inline at no separate caption fee. Add product_image if the host should show the item; it adds $0.013 per second at plus.
When the video model wins
A reference-driven ad that reuses the framing of a winning clip needs references, not a script. Sume's video catalog accepts image references on all reference-capable models and audio and video references on the Seedance 2.x models (check the model's entry in GET /v1/videos/models). Pin the model id seedance-2.5 on POST /v1/videos, or use sume/auto to let Sume select one.
Decision
Count the speakers. One speaker plus a script is an avatar job; zero speakers or several locations is a video job. For both in one ad, make the talking hook as an avatar clip and the b-roll as a video job, and cut them in a Timeline render.
Worked cost for 30 seconds
Avatar 1.0 charges per second of final video: $0.184 standard, $0.245 plus, $0.55 max with no product image. Thirty seconds is therefore $5.52, $7.35 or $16.50. The seedance-2.5 price is not quoted here because it belongs in the live catalog: GET /v1/videos/models returns pricing_skus for each model, so read it before you budget a batch.
- Avatar job: one handle, one script, one aspect ratio.
- Video job: prompt plus optional references, one clip up to 30 seconds.
Mixing both in one ad
Many 30-second ads do not need to pick. Render the talking hook as a short Avatar 1.0 job, render the b-roll as a video job, and join them in a Timeline 1.0 render at $0.10 per output minute. That keeps the face and the words exact where they matter and lets the video model carry the motion. Plan the join first with the unbilled plan endpoint so you know the total before paying.
Sources
Related posts
More in Use cases
- 4:3 presentation background: 10 seconds on Seedance 2.5, $5.82
A 10-second 4:3 slide background at 720p on Seedance 2.5 renders 1112x834 and costs $5.82 on Sume. Why 4:3 beats cropping 16:9, with a loop tip.
- 4:5 AI video: HappyHorse lists it, Sume's rows do not. Crop instead.
Alibaba's HappyHorse text-to-video reference lists 4:5, 5:4 and 9:21. None of Sume's H3, Kling or Omni rows takes 4:5, so generate 3:4 or 1:1 and crop.
- 40 micro-drama episodes of 60 seconds on Wan 3.0: $150 at 480p
Forty 60-second micro-drama episodes are two 30-second Wan 3.0 jobs each: $150.00 at 480p, $300.00 at 720p and $600.00 at 1080p on Sume.
- 5-minute AI presenter video on Sume: $55.20 avatar plus $0.50 render
A 5-minute presenter video made of five 60-second Standard avatar jobs (no product) costs $55.20, plus $0.50 to render the Timeline: $55.70 in total.
Written by Sume