Azure avatar batch: 20-minute videos, 200 jobs, vs Sume

Azure's text to speech avatar batch API allows 20-minute outputs and 200 concurrent jobs per Speech resource. Sume avatar videos run 4-60 seconds per job.

4 min readSume
All posts

Azure's text to speech avatar batch API caps one output video at 20 minutes and allows up to 200 batch jobs running at once per Speech resource. Sume's avatar video is a shorter unit: a script or scene plan estimated at 4-60 seconds per job, with concurrency set by your plan.

Both are asynchronous: submit, poll, download. The limits below come from Microsoft's batch synthesis page and Sume's avatar video docs.

What are Azure's batch limits?

The page says each Speech resource can have up to 200 batch synthesis jobs running concurrently, the maximum JSON payload is 500 kilobytes, and the maximum output video is currently 20 minutes, with potential increases in the future. Input is plain text or SSML, and the SynthesisId is 3 to 64 characters.

A job moves from NotStarted to Running and finally to Succeeded or Failed. You list jobs with skip and maxpagesize, whose default and maximum page size is 100.

What are Sume's?

An avatar video accepts a script, or ordered video_inputs, when Sume estimates 4-60 seconds inclusive. Longer scripts must be shortened or split into several jobs. A script holds one avatar per final video, and video_inputs can have scene hooks, demos and silence beats.

Concurrency is a plan limit rather than a per-resource constant: Free is 1 processing job with a queue of 5, Pro is 4 with 20, Startup 8 with 40, and Scale 20 with 100, so accepted job capacity for Pro is 24.

Azure from the Microsoft Learn batch synthesis page; Sume from Generate avatar video and Generation admission; read 2026-10-03.
LimitAzure text to speech avatar batchSume avatar video
Max length of one video20 minutes60 seconds estimated, 4 minimum
Concurrent work200 batch jobs per Speech resourcePlan-based: Free 1, Pro 4, Startup 8, Scale 20 processing
Over capacityNot described on the pageValid jobs queue; 429 queue_full when the queue is full
Status valuesNotStarted, Running, Succeeded, Failedqueued, processing, completed, failed, canceled
Text inputPlain text or SSMLscript or video_inputs

How do I make a long video on Sume?

Split it. A ten-minute lesson is a set of scene-sized jobs of up to 60 seconds, each submitted with its own Idempotency-Key and the same avatar_handle, then joined. Sume's Timeline 1.0 is the assembly surface for that: an audio spine plus ordered video clips into one MP4.

How long do the results last?

Azure keeps each synthesis history for up to 31 days or the timeToLiveInHours you set, whichever comes sooner, and the output link carries a SAS token. Download the file inside that window.

Sume results are public media.sume.com artifacts on the finished job, and you read them from GET /v1/jobs/{id}/result. Read the lifetime of any file you depend on before building a retention plan on top of it.

Which limit should drive the choice?

If you need one long unbroken presenter video, the 20-minute ceiling matters more than anything else on the page. If you make many short clips and want a plan-based queue that accepts work without you tracking an upper bound, Sume's admission model is the one to size against.

What would a 5-minute video cost in jobs?

At the 60-second ceiling, a five-minute script is at least five Sume jobs, and in practice more because you want each scene to end on a sentence. Plan the cut points in the script itself, keep one avatar_handle across all of them, and submit each with its own Idempotency-Key so a retry never double-bills.

On a Pro workspace, whose accepted job capacity is 24, all of those jobs can be accepted at once; on Free, with a capacity of 6, a larger batch gets 429 queue_full and you wait for jobs to finish. The generation_limits block in each submit response tells you which situation you are in.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume