Moving Company Spokesperson Ad: A 20-Second Avatar for About $4.90
A 20-second plus-quality Avatar 1.0 spokesperson clip costs about $4.90 on Sume, at 98 cents per 4 seconds. English only, with captions at $0.20 extra.
What does a spokesperson video for a moving company cost with an AI avatar? A 20-second talking clip at plus quality is about $4.90 on Sume Avatar 1.0, because the golden test prices 4 seconds at 98 cents and 20 seconds is five of those. Add $0.20 if you want captions burned in, for $5.10.
One limit decides whether this fits you: Avatar 1.0 is English only. If your customers speak another language, this post does not solve it.
The script that fits 20 seconds
Twenty seconds is roughly 50 spoken words. A moving ad has one job, so write exactly three beats: who you are and where you work, the one thing that makes you different, and what the caller should do. For example, that is a local company name, a fixed-quote promise and a phone call to action.
The Generate avatar video route takes an avatar_handle plus either a script or video_inputs, never both. The estimated duration must be 4 to 60 seconds, so 20 seconds is comfortably inside the range.
Quality tiers and what they cost
Quality is standard, plus (the default) or max. The golden prices for 4 seconds are 74 cents, 98 cents and 220 cents. Scaled to 20 seconds, that is $3.68 for standard, $4.90 for plus and $11.00 for max. Standard is the cheapest and fastest tier; plus is the default.
Aspect ratio defaults to 9:16, which suits a vertical ad. Other options are 1:1, 3:4, 4:3 and 16:9, and the output resolution is 720p.
| Quality | 4 seconds | 20 seconds | With $0.20 captions |
|---|---|---|---|
| standard | 74 cents | $3.68 | $3.88 |
| plus (default) | 98 cents | $4.90 | $5.10 |
| max | 220 cents | $11.00 | $11.20 |
Captions, background and the truck
Inline captions are an option on the avatar request itself. The default style is slam, and there are punch and tiktok-green styles too. In the golden test, adding them took the 4-second plus clip from 98 to 118 cents. The docs say caption styling fails soft, so a caption problem does not cost you the video.
An optional product_image and scene let you put something in the shot. For a mover, that could be a clean photo of your truck. Use a real photo, not a generated one, because customers will look for that truck on moving day. Adding a product image raises the plus rate in the golden prices from $0.245 to $0.258 per second, so the 20-second clip would be $5.16.
Set expectations honestly
An avatar speaks as a presenter, not as your actual foreman. Say so if a customer asks, and never give the avatar a credential or a review quote you cannot back up. Voice cloning is only available in the app and not through the API, so do not plan on your own cloned voice for an API-made clip.
Test with the 4-second minimum at standard quality first. That costs 74 cents and tells you whether the look and the voice suit the brand before you spend $4.90 on the real script.
Previewing before you pay for the full clip
The docs describe avatar video previews: you request a preview, can regenerate it, and then generate the video. A preview is the first-frame still stage and does not start the full render, so you can approve the composition before paying for the full video.
Plan the run of ads as well. If you want a spring and a summer version with different offers, change only the script and reuse the avatar handle, so every ad looks like the same spokesperson. Ten 20-second plus clips would cost 10 x $4.90 = $49.00, and the consistency is worth more than the novelty of a new face each time.
Remember the 60-second ceiling. A longer story should be split into several clips, each with its own script, and joined later with a timeline render.
Think about where the clip will run. A vertical 9:16 ad suits a social feed, while a 16:9 version fits a website header. Both are listed aspect ratios, and each is a separate job, so decide the main placement first and generate the second only if the first performs. Keep the phone number large in the caption text and say it twice, once in speech and once on screen, so a viewer can act on either.
Sources
Related posts
More in Sume Avatar 1.0
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
- Avatar video previews: approve the first frame before rendering
Create an avatar video preview to get first-frame stills, regenerate them if needed, then call generate-video on the preview id to render the final video.
- How to create a reusable AI avatar with the Sume Avatar 1.0 API
Send POST /v1/avatar-1.0/generate with an avatar_handle and a prompt, profile, or image input. Poll the job, then reuse the handle for avatar videos.
Written by Sume