Azure avatar batch: 20-minute videos, 200 jobs, vs Sume
Azure's text to speech avatar batch API allows 20-minute outputs and 200 concurrent jobs per Speech resource. Sume avatar videos run 4-60 seconds per job.
Azure's text to speech avatar batch API caps one output video at 20 minutes and allows up to 200 batch jobs running at once per Speech resource. Sume's avatar video is a shorter unit: a script or scene plan estimated at 4-60 seconds per job, with concurrency set by your plan.
Both are asynchronous: submit, poll, download. The limits below come from Microsoft's batch synthesis page and Sume's avatar video docs.
What are Azure's batch limits?
The page says each Speech resource can have up to 200 batch synthesis jobs running concurrently, the maximum JSON payload is 500 kilobytes, and the maximum output video is currently 20 minutes, with potential increases in the future. Input is plain text or SSML, and the SynthesisId is 3 to 64 characters.
A job moves from NotStarted to Running and finally to Succeeded or Failed. You list jobs with skip and maxpagesize, whose default and maximum page size is 100.
What are Sume's?
An avatar video accepts a script, or ordered video_inputs, when Sume estimates 4-60 seconds inclusive. Longer scripts must be shortened or split into several jobs. A script holds one avatar per final video, and video_inputs can have scene hooks, demos and silence beats.
Concurrency is a plan limit rather than a per-resource constant: Free is 1 processing job with a queue of 5, Pro is 4 with 20, Startup 8 with 40, and Scale 20 with 100, so accepted job capacity for Pro is 24.
| Limit | Azure text to speech avatar batch | Sume avatar video |
|---|---|---|
| Max length of one video | 20 minutes | 60 seconds estimated, 4 minimum |
| Concurrent work | 200 batch jobs per Speech resource | Plan-based: Free 1, Pro 4, Startup 8, Scale 20 processing |
| Over capacity | Not described on the page | Valid jobs queue; 429 queue_full when the queue is full |
| Status values | NotStarted, Running, Succeeded, Failed | queued, processing, completed, failed, canceled |
| Text input | Plain text or SSML | script or video_inputs |
How do I make a long video on Sume?
Split it. A ten-minute lesson is a set of scene-sized jobs of up to 60 seconds, each submitted with its own Idempotency-Key and the same avatar_handle, then joined. Sume's Timeline 1.0 is the assembly surface for that: an audio spine plus ordered video clips into one MP4.
How long do the results last?
Azure keeps each synthesis history for up to 31 days or the timeToLiveInHours you set, whichever comes sooner, and the output link carries a SAS token. Download the file inside that window.
Sume results are public media.sume.com artifacts on the finished job, and you read them from GET /v1/jobs/{id}/result. Read the lifetime of any file you depend on before building a retention plan on top of it.
Which limit should drive the choice?
If you need one long unbroken presenter video, the 20-minute ceiling matters more than anything else on the page. If you make many short clips and want a plan-based queue that accepts work without you tracking an upper bound, Sume's admission model is the one to size against.
What would a 5-minute video cost in jobs?
At the 60-second ceiling, a five-minute script is at least five Sume jobs, and in practice more because you want each scene to end on a sentence. Plan the cut points in the script itself, keep one avatar_handle across all of them, and submit each with its own Idempotency-Key so a retry never double-bills.
On a Pro workspace, whose accepted job capacity is 24, all of those jobs can be accepted at once; on Free, with a capacity of 6, a larger batch gets 429 queue_full and you wait for jobs to finish. The generation_limits block in each submit response tells you which situation you are in.
Sources
Related posts
More in Sume Avatar 1.0
- How to check a face swap result: audio, length and stills
Check a Sume face swap before publishing: confirm the audio stream, compare length with the source, and sample stills with video inspect. Free probe.
- Create an AI spokesperson avatar from a prompt or photo
Sume Avatar 1.0 creates a reusable avatar at a flat $0.95 from a text prompt, a photo, or props. How to make one before you generate talking videos.
- Demand Gen: what to hold constant between two AI video ads
Google says Demand Gen experiment campaigns should differ in one variable. A field-by-field checklist for two Sume avatar videos so only the hook changes.
- Avatar face swap Beta: check the 4 to 15 second clip before you submit
Sume's Avatar Face Swap Beta needs a public HTTPS source clip of about 4 to 15 seconds with usable audio and a required quality tier.
Written by Sume