Azure photo avatar is 512x512; Sume avatar video is 720p
Azure's photo avatar renders head-only at 512x512 and 25 fps. Sume's avatar video offers 720p in five aspect ratios, from a prompt, traits or a photo.
Azure's photo avatar is head-only and renders at 512x512, 25 frames per second, in both batch and real-time synthesis. Sume's avatar video renders at 720p today, in 1:1, 3:4, 9:16, 4:3 or 16:9, so a vertical clip is a real vertical frame.
The two are not the same kind of product, but if you are picking a photo-to-talking-head API by output specs, those two numbers are the first thing to compare. The Azure facts are from Microsoft's text to speech avatar overview, updated August 18, 2026.
What does Azure's photo avatar produce?
Photo Avatar is created from a single input image, is limited to a head-only representation, and comes in a standard variant and a custom variant that lets you fine-tune the look from your own image. Microsoft says it uses its VASA model. Photo avatar resolution is 512x512 for batch and real-time, at 25 fps. Batch codecs are H264, HEVC or VP9.
A custom photo avatar needs a photo plus about one minute of consent audio, which the service uses to match the voice to the avatar. Video avatars, a different type, default to 1920 x 1080 and can be trained for 4K.
What does Sume produce?
resolution is currently 720p. aspect_ratio accepts 1:1, 3:4, 9:16, 4:3 and 16:9, with 9:16 as the default. quality is standard, plus (default) or max, which changes the execution path, not the resolution.
The avatar can be created from a prompt, structured props, or a reference photo, and you then reuse its handle. A script of an estimated 4-60 seconds produces one MP4.
| Spec | Azure photo avatar | Sume avatar video |
|---|---|---|
| Resolution | 512x512 | 720p |
| Aspect | Square | 1:1, 3:4, 9:16, 4:3, 16:9 |
| Frame rate | 25 fps | Not stated in the docs |
| Body framing | Head only | Not specified in the docs |
| Source | A single photo (custom adds about 1 minute of consent audio) | Prompt, traits or a photo |
Does Sume have a 1080p or 4K option?
Not in the docs: Sume lists 720p as the only current resolution. If your deliverable must be 1080p, test a clip first and check it with a probe before you commit a batch. Sume's earlier post on 4K and 1080p avatar video covers the upscale question in detail.
How should I choose?
For a 512x512 head in a square avatar widget, Azure's photo avatar matches the use. For a 9:16 social clip, a presenter frame, or a landing-page video, the aspect ratio list matters more than the pixel count, and Sume's 720p vertical is usually the closer fit.
Whichever you use, ask for a short sample and view it on the target device. Specs tell you the frame, not whether the mouth looks right on your face.
What else differs between the two?
Azure bills the text to speech separately from the avatar, and says avatar pricing is visible only for supported regions, so compare the combined cost, not the avatar line alone. Sume reserves the cost of the job at submit, and the plan queue limits apply to avatar jobs like any other paid generation.
If your use is a head in a small widget, 512x512 may be all you need. If the clip will fill a phone screen, look at the 9:16 default and the preview stills before you decide that frame size is the whole story.
Sources
Related posts
More in Sume Avatar 1.0
- Azure avatar batch: 20-minute videos, 200 jobs, vs Sume
Azure's text to speech avatar batch API allows 20-minute outputs and 200 concurrent jobs per Speech resource. Sume avatar videos run 4-60 seconds per job.
- How to check a face swap result: audio, length and stills
Check a Sume face swap before publishing: confirm the audio stream, compare length with the source, and sample stills with video inspect. Free probe.
- Create an AI spokesperson avatar from a prompt or photo
Sume Avatar 1.0 creates a reusable avatar at a flat $0.95 from a text prompt, a photo, or props. How to make one before you generate talking videos.
- Demand Gen: what to hold constant between two AI video ads
Google says Demand Gen experiment campaigns should differ in one variable. A field-by-field checklist for two Sume avatar videos so only the hook changes.
Written by Sume