10-minute avatar video from 60-second jobs: HeyGen cap vs Sume

HeyGen Creator renders up to 30 minutes. Sume avatar jobs run 4-60 seconds, so a 10-minute course is ten jobs joined by a $1.00 Timeline render.

5 min readSume
All posts

For a 10-minute talking-avatar course, HeyGen Creator does it in one video and Sume does it as at least ten jobs joined afterward. HeyGen's pricing page lists 30 minutes per video on Creator and Pro, and 60 on Business; Synthesia Basic lists 10 minutes of video a month. Sume accepts avatar scripts only when the estimated duration is 4 to 60 seconds.

The join step is cheap. Timeline 1.0 is $0.10 per output minute, so a 10-minute course is $1.00 to assemble. The cost to plan for is the ten avatar jobs, whose price comes from the live catalog.

Length limits for a 10-minute avatar video, read 2026-10-08
ToolLimit that matters10-minute course
HeyGen Creator ($29)30 minutes per video; 600 credits a monthOne video
Synthesia Basic (free)10 minutes of video a month; 500 creditsUses the whole month's allowance
Sume Avatar 1.04 to 60 seconds per job10 jobs of 60 seconds
Sume Timeline 1.0audio.duration_seconds 1 to 1,800; up to 200 video slots600 seconds = 10 billable minutes = $1.00

How to build it on Sume

Write the course as ten scripts of about a minute each, and call POST /v1/avatar-1.0/talking-video ten times with the same avatar_handle so the presenter is stable. Poll each job, or use a webhook. Each result is a Sume-hosted video.

Then assemble. The simplest route is to place each clip as a slot in Timeline 1.0, with the clip audio detached and joined into one spine through Timeline audio ($0.01 per job). Timeline audio can concatenate hosted audio, and audio.parts[] accepts up to 20 gapless slices if you prefer to do it inside the render.

  • Avatar: 1:1, 3:4, 9:16, 4:3 or 16:9, 720p.
  • Timeline: fade, wipe, slide or dissolve transitions up to 1 second; default output 1080 by 1920 unless you set width and height.
  • Plan first: POST /v1/timeline-1.0/plan returns billable_minutes and the estimate without a charge.

Seams to watch

Ten separately generated clips can differ slightly in lighting or framing. Use avatar video previews to approve the first frame of each scene before rendering, or send multi-scene video_inputs if a single 60-second job can cover a chapter.

Inline captions are accepted only when the estimated duration is 60 seconds or less, so caption each chapter clip before joining, or caption the final file with a caption job, which is $0.20 for videos up to 60 seconds. A 10-minute file is outside that fixed estimate, so caption the pieces.

A checklist for the build keeps the pieces in order. Define the avatar once and reuse its handle. Write the scripts so that each one ends on a clean sentence, because a hard cut in the middle of a thought is more noticeable between separate clips. Generate previews for the first and last chapters at least. Detach, join and render the audio and video in Timeline, using the plan call to check that the total is 10 billable minutes. Finally, caption the chapter clips or the final file as your length allows, and store the job ids next to the project so that a single chapter can be regenerated without redoing the course.

Verdict

If the content is one long lecture, use a tool whose page lists long videos. If it is a series of short lessons, or an ad plus variations that you submit from code, the 60-second window is not a restriction and the join costs about a dollar.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume