Course sale avatar clip in five aspect ratios: five Sume calls

One course-sale script, five aspect ratios: avatar talking-video takes 1:1, 3:4, 9:16, 4:3 and 16:9 per call at 720p, 4 to 60 seconds. Plan five calls, not one.

4 min readSume
All posts

If you sell a course in the holiday weeks and want the same avatar message in every placement, you make five separate calls to POST /v1/avatar-1.0/talking-video, one per aspect_ratio. The endpoint accepts 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 as the default, and the resolution is 720p at this time. A single request produces one video in one ratio; it does not return a set.

The facts to plan around

The avatar doc fixes the total length at 4 to 60 seconds for a script or for the sum of the scenes in video_inputs, and you send either script or video_inputs, never both. The avatar_handle names a ready avatar.

Avatar talking-video parameters for a multi-ratio sale clip, read 2026-10-08
ParameterValues or rule
aspect_ratio1:1, 3:4, 9:16 (default), 4:3, 16:9
resolution720p at this time
Total duration4 to 60 seconds
qualitystandard, plus (default), max
Script inputscript or video_inputs, not both
IdentityTop-level avatar_handle
Retry safetyIdempotency-Key header

Steps

Write once, then vary only what each placement needs.

  • Keep one script of the length you want, well inside the 4 to 60 second window, so each ratio has the same story.
  • Loop over the five ratios and send one request each, with a distinct Idempotency-Key such as course-sale-16x9, so a network retry returns the same job and not a second paid one.
  • Keep quality the same across the five calls unless you are testing, since max is the slower tier and standard is the fastest path.
  • Poll each job, then place the finished files by ratio: square for feeds, 9:16 for stories and short video, 16:9 for a landing page.

What Sume does not do

Sume does not reframe one finished video into the other four ratios; each ratio is a fresh generation, and the framing of the avatar can differ between them. Look at all five before you publish. Resolution stays at 720p, so a large-screen placement is not a 4K asset.

The docs also say Sume does not add captions to the video by itself. If you want burned-in text, request inline captions in the call or run a separate caption job, and plan for the extra cost shown in the live catalog rather than assuming it is included.

A sensible sale-week routine is to generate the 9:16 version first, because it is the default and the most common placement, review it, and only then spend on the other four. If the script needs a change after review, you have paid for one video and not five. Once the script is final, submit the remaining four ratios together with their own idempotency keys, and keep the keys in your launch checklist so a re-run of the checklist cannot buy the same clip again.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume