AI avatar video aspect ratio: square, 3:4 and vertical 9:16
Sume's avatar video endpoint takes aspect_ratio 1:1, 3:4, 9:16, 4:3 or 16:9, defaulting to 9:16, at 720p. Which value to send for square or vertical.
Sume's avatar video endpoint supports five aspect ratios: 1:1, 3:4, 9:16, 4:3 and 16:9. Send aspect_ratio: "1:1" for square and "9:16" for vertical; if you omit the field you get 9:16. Resolution is currently 720p for every ratio.
Which aspect ratio value do I send?
The field goes on POST /v1/avatar-1.0/talking-video, next to avatar_handle and either script or video_inputs. Match the value to the frame you plan to place the video in.
| aspect_ratio | Shape | Note |
|---|---|---|
9:16 | Vertical | The default when omitted. |
3:4 | Tall portrait | Less tall than 9:16. |
1:1 | Square | Equal width and height. |
4:3 | Landscape | The wider counterpart of 3:4. |
16:9 | Wide landscape | Widescreen frame. |
Does the ratio change the resolution?
No. resolution is currently 720p and the docs list no other value, so the ratio changes the shape of the frame, not the resolution setting. This page does not state any platform's own size requirements; check the destination's help page for those before you pick a ratio. A ratio the endpoint accepts is not a promise that a placement will use it unchanged, so preview a first frame in the shape you chose.
How do I set it on a multi-scene video?
Set aspect_ratio once at the top level, alongside the ordered video_inputs; the docs example sends "aspect_ratio": "9:16" with three scenes. Current execution supports one resolved avatar per final video and expects scene backgrounds to resolve to one shared scene, so the ratio applies to the whole video rather than per scene. If you need the same script in both a square and a vertical placement, send two requests that differ only in aspect_ratio, each with its own Idempotency-Key, since a reused key with a different body returns 409 idempotency_conflict.
Can I check the framing before I pay for the render?
Yes. To review first-frame stills before paying for a full render, create an avatar video preview, then call generate-video on the preview id. See avatar video first-frame previews. Full parameters are in the avatar video docs.
What else should I set alongside the ratio?
Avatar Video accepts quality of standard, plus or max. plus is the default when omitted, standard is the quick Sume execution path, and max is the highest quality tier with slower turnaround. Quality is a separate field from the ratio.
Scripts and multi-scene plans are accepted when Sume estimates the target video duration at 4 to 60 seconds inclusive. Shorten longer scripts or split them into multiple jobs.
Optional inline captions burn styles into the clean final MP4 after generation, so decide the ratio before you enable them.
Sources
Related posts
More in Models
- AI image negative prompt: what to do when the API has no field
Sume's image API has no negative prompt field. Describe what you want instead, edit with a mask, and keep references. The fields that do exist.
- AI image generator with readable text: which Sume models to try
Which Sume image models have a vendor claim about text in images, what each vendor says, and the quality setting Sume's docs suggest for dense text.
- Does Lyria 3.5 music carry a SynthID watermark?
Google says all Lyria 3.5 audio includes a SynthID watermark. What that means for tracks made through an API, and what Sume's docs do and do not say.
- AI sky replacement: swap the sky and keep the photo
AI sky replacement: send the photo to an image-edit model, describe only the new sky, and list what stays. Then check roof and tree edges. Cost and limits.
Written by Sume