3-minute AI presenter video: three 60-second clips plus Timeline
A 180-second presenter video built from three 60-second avatar jobs costs about $44.10 on Plus plus $0.30 to join them with Timeline, at listed rates.
A single Sume avatar job tops out at 60 seconds, so a 3-minute presenter video is three jobs joined into one MP4. At the listed Plus rate that is $44.10 of video, plus about $0.30 to render the join with Timeline (a per-output-minute charge, read 2026-10-03), for roughly $44.40 before the one-time avatar charge.
The 60-second ceiling is documented on the avatar video page: a script or multi-scene plan must estimate at 4 to 60 seconds, and longer scripts should be shortened or split into several jobs. Nothing in the avatar docs stitches the parts for you, which is where Timeline comes in.
Split on a thought boundary
Cut the script where a viewer would naturally take a breath: after the hook and promise, after the main demonstration, before the call to action. Each part should open in the same framing so the join does not jump. Using one avatar_handle and the same background prompt for all three jobs keeps the person and setting consistent.
Keep every part comfortably under the limit. A script that Sume estimates at 61 seconds is rejected, so aim for 55 to 58 seconds of speech per part and leave room.
What it costs
The Timeline render is priced per output minute; a 3-minute output is three of them. Timeline also requires that every input URL is already a Sume-hosted artifact in your workspace, which avatar results are, because the completed job returns media.sume.com video URLs (see the Timeline docs).
| Tier | 3 x 60 s video | Timeline join | Total |
|---|---|---|---|
| Standard | $33.12 | $0.30 | $33.42 |
| Plus | $44.10 | $0.30 | $44.40 |
| Max | $99.00 | $0.30 | $99.30 |
Estimate in code
Timeline offers an unbilled plan route that returns an estimate before you commit, so use it for the join. For the avatar parts the arithmetic below is enough.
import asyncio
def clip_cost(seconds: float, rate: float = 0.245) -> float:
return seconds * rate
async def main() -> None:
parts = [60, 60, 60] # three clips, each inside the 4-60 s window
video = sum(clip_cost(s) for s in parts)
join = 0.10 * 3 # timeline render, per output minute
print(f"video ${video:.2f} + timeline ${join:.2f} = ${video + join:.2f}")
asyncio.run(main())
When not to join
If the video is a series of separate lessons, do not join at all: publish three 60-second videos. Joining only makes sense when the viewer should experience one continuous piece, such as a 3-minute product walkthrough. For ideas on keeping scenes varied inside one part, see adding a silence beat.
Keeping the person consistent across parts
Three separate jobs mean three separate renders, so consistency has to come from your inputs. Use the same avatar_handle, the same aspect_ratio and quality, and the same background prompt text in every part. The docs say that current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, which is the natural fit for a part that repeats a setting.
Avoid asking part two to change the room or the outfit. If you need a location change, treat it as a deliberate cut and plan the transition in the Timeline document rather than hoping two independent renders will match.
Where the join goes wrong
None of this is difficult, but it is more moving parts than a single job. If a 60-second video can carry your message, it is cheaper, simpler and usually better for completion rates.
- Audio spine: Timeline takes one audio spine plus ordered video slots, so you must supply the audio for the joined video. Detach each clip's audio and concatenate it in a prior step, or follow the Timeline audio route.
- Hosted media only: every URL must be a media.sume.com artifact in your workspace, so import any outside footage first.
- Length: Timeline allows up to 1800 seconds of output, so three minutes is far inside the limit.
- Preview cost: the plan route is unbilled and returns the estimated cost and segment count before you render.
Sources
Related posts
More in Use cases
- 3-minute explainer Reel as six beats: voice first, then pictures
Plan a 180-second Reel as six voiced beats: one TTS file per beat, one concat for offsets, then picture slots that start where each beat starts.
- 3-minute Reel script with AI voice: characters, cost and the cap
A 180-second Reel narration is a character budget, not a word count. Sume TTS is $0.0475 per 1,000 characters with a 20,000-character cap; measure, then scale.
- A/B test three Shorts hooks from one body: Timeline source_in
YouTube announced A/B testing for up to three Short versions. Build the three files from one body with Timeline slots, source_in and a plan call before you pay.
- Add a related video to a YouTube Short made outside YouTube
A related video is a clickable link under a Short's channel name. YouTube says it needs advanced features and a public or unlisted video. Set it in Studio.
Written by Sume