Reuse one AI presenter across lip-sync clips with an avatar handle
Create the presenter once, then pass avatar_handle to H3 Max lip sync for every clip. Same face each time, no re-upload, one Idempotency-Key per line.
Create the presenter once with POST /v1/avatar-1.0/generate, keep its handle, and pass avatar_handle to the H3 Max lip-sync route for every clip. The face comes from the avatar, so you never re-upload a photo and the presenter does not drift between episodes. The lip-sync body takes exactly one of image_url, avatar_id or avatar_handle, plus a Sume-hosted audio_url of 5 to 14.8 seconds and a duration_seconds in the same range.
Routes and limits here come from Models, Create new avatar and Generate avatar video, read 2026-10-06.
Why a handle instead of an image URL?
| Input | Face source | Best for | Watch for |
|---|---|---|---|
| image_url | A still you host | A one-off face | Aspect ratio must be 0.4 to 2.5 |
| avatar_id | A stored avatar | Programmatic reuse | Id must exist in your account |
| avatar_handle | A stored avatar by name | Series and teams | Handle rules, no leading @ stored |
What does a batch look like?
Swap the handle and audio URLs for yours. Each line gets its own Idempotency-Key, so a retry cannot double-charge, and each reserves its own ceil(duration) times the rate: 7 s plus 11 s at 768p is $0.70 plus $1.10.
import asyncio, os, httpx
URL = "https://api.sume.com/v1/minimax/h3-max/lip-sync"
KEY = os.environ["SUME_API_KEY"]
LINES = [("intro", "https://media.sume.com/demo/intro.mp3", 7),
("offer", "https://media.sume.com/demo/offer.mp3", 11)]
async def one(c, name, audio, secs):
body = {"avatar_handle": "studio-host", "audio_url": audio,
"duration_seconds": secs, "resolution": "768p"}
h = {"Authorization": "Bearer " + KEY,
"Idempotency-Key": "host-" + name}
r = await c.post(URL, json=body, headers=h)
return name, r.status_code
async def main():
async with httpx.AsyncClient(timeout=30) as c:
out = await asyncio.gather(*[one(c, *x) for x in LINES])
print(out)
asyncio.run(main())
What the handle must look like
A handle is 2 to 30 characters of lowercase letters, numbers, underscores or periods, and periods and underscores cannot be first, last or consecutive. Create the avatar with a leading @ if you like: Sume normalizes the handle and stores it without the @, so pass the plain form to the lip-sync route. The API validates the shape at submit, which means a bad name fails before anything is created or charged (Create new avatar).
The three face inputs on the lip-sync route are exclusive: image_url on one side, and avatar_id or avatar_handle on the other. If you send an image_url, it must be a public HTTPS still with an aspect ratio between 0.4 and 2.5. The audio goes through the same Sume-host and 10 MB preflight as the other talking-video routes, and the errors unsupported_audio_source and audio_too_large tell you which check failed.
Two routes reach the same model: POST /v1/minimax/h3-max/lip-sync and the canonical model-run twin at POST /v1/models/minimax/h3-max/lip-sync/runs. They store the same public model id, so use whichever fits the client you already have, and do not mix them up with the prompt-driven Video Router model minimax-h3-max, which is a different product.
What limits a batch?
Plan concurrency decides how many run at once: Free 1, Pro 4, Startup 8, Scale 20, with queue room of 5, 20, 40 and 100 behind them. Beyond that you receive a 429 queue_full, and a balance too low for the reserve returns 402 insufficient_credits before any work starts. Stagger submissions to your plan rather than firing everything at once, and poll or use webhooks for completion.
If the avatar has no usable voice, Sume's talking-video route cannot speak for it, which is a 400; lip sync avoids this because you supply the audio. See the voice error post and the handle rules before you name the presenter, and the mascot post for a brand version.
How do you keep the presenter consistent over months?
Treat the handle like a production asset. Write down who created it, the source photo or prompt, and the date. Do not delete or rename it while episodes depend on it, and never create a second avatar for the same person to fix a bad render; regenerate the clip instead. A handle that already exists returns 409 avatar_handle_taken, and one that starts with the system-only sume_ prefix returns avatar_handle_reserved, so check naming before the first episode.
Add a consent record for the person whose likeness the avatar uses. Reusing a face across many clips makes that record more important, not less.
Sources
Related posts
More in Sume Avatar 1.0
- Tavus Griffin-Lite is invite-only: what can you build today?
Griffin-Lite is a closed Tavus preview. Until you get in, Sume ships rendered avatar clips, lip sync and motion control for talking-head video.
- YouTube Shorts series: a weekly AI presenter episode pipeline
YouTube began rolling out Shorts series on 2026-09-23. A repeatable weekly pipeline for presenter episodes: one avatar, draft at standard, final at max.
- Introducing Sume Avatar 1.0
Sume Avatar 1.0 is a multi-agent orchestration system as a single avatar model.
- Avatar Face Swap API (Beta): apply an avatar face to a video
Avatar Face Swap 1.0 is a Beta Sume endpoint that applies a ready avatar's face to a short public source video. Required fields, limits, and polling.
Written by Sume