3-minute AI presenter Reel: three 60-second avatar jobs, cost by tier
An avatar job accepts 4 to 60 seconds, so a 180-second presenter Reel is three jobs: $33.12 on Standard, $44.10 on Plus, $99.00 on Max at no-product rates.

A talking-avatar job on Sume accepts a script that estimates between 4 and 60 seconds, so a 3-minute presenter Reel is three jobs of up to 60 seconds. At the no-product per-second rates, 180 seconds of avatar video costs $33.12 on Standard, $44.10 on Plus and $99.00 on Max, before assembly: $0.01 per audio detach (three of them) and the $0.30 Timeline render.
The arithmetic
Rates come from the catalog rate card for sume/avatar-1.0/talking-video (read 2026-10-03): Standard $0.184 per second, Plus $0.245 and Max $0.55, with product-image variants slightly higher. Output is 720p, and the default aspect ratio is 9:16. Standard is the fastest tier and Max the highest quality and slower; Plus is the default.
| Tier | Per second | One 60 s job | Three jobs (180 s) |
|---|---|---|---|
| standard | $0.184 | $11.04 | $33.12 |
| plus | $0.245 | $14.70 | $44.10 |
| max | $0.55 | $33.00 | $99.00 |
Plan the split
Metricool (read 2026-10-03) reports a 3-minute Reel length, and Instagram's guide (read 2026-10-03) lists 60 seconds to 3 minutes for longer storytelling. Split the script at sentence ends into three parts of about equal estimated length, so no part exceeds 60 seconds; a part that estimates over 60 is rejected, and one under 4 is billed at 4.
import re
def split_script(text, parts=3):
sents = re.split(r"(?<=[.!?])\s+", text.strip())
target = len(text) / parts
out, cur = [], ""
for s in sents:
if cur and len(cur) + len(s) > target and len(out) < parts - 1:
out.append(cur.strip())
cur = ""
cur += s + " "
out.append(cur.strip())
return out
for p in split_script("One. Two is here. Three follows. Four ends. " * 6):
print(len(p))Join the three clips
Put the three MP4s on consecutive slots of Timeline 1.0. The render takes its sound from one audio spine, so detach each clip's voice ($0.01 each) and join the three files with audio.parts[] in the render, or into one file with timeline audio ($0.01). Use a hard cut at each join, since the presenter returns to a similar pose, or add a short fade (up to 1 second). Preview the first part with the preview flow before you commit to three paid renders.
What Sume does not do
Sume does not keep gesture continuity between separate jobs and does not provide a single 180-second presenter job. Confirm the live prices with GET /v1/catalog before a batch; the figures here are dated.
Sources
Related posts
More in Sume Avatar 1.0
- AI avatar video with a transparent background: Azure webm vs Sume
Azure's avatar batch API can output VP9 webm with an alpha channel. Sume's avatar video returns an MP4 with a scene behind the avatar.
- AI spokesperson: build the avatar from a prompt, profile or photo?
Sume makes an avatar three ways: prompt, profile traits or a photo. Which one suits a brand spokesperson, what each costs, and the rights each one raises.
- AI spokesperson video price per second: standard, plus, max
Sume Avatar 1.0 talking video bills per second: $0.184 standard, $0.245 plus, $0.55 max, with a product image adding a little. A 4-second floor example.
- Avatar 1.0 API routes: canonical paths vs legacy aliases
New Avatar 1.0 code should call /v1/avatar-1.0/generate and /v1/avatar-1.0/talking-video. Older model-run and legacy aliases still work with the same body.
Written by Sume