Overlay a graphic on avatar video: Sume compose overlay step
HeyGen's demo fades a HyperFrames composition over a live avatar. On Sume, put one still over finished avatar footage with timeline compose operation overlay.
On Sume you overlay a graphic on avatar footage as a second step: render the avatar video, import it and a still plate, then call POST /v1/timeline-1.0/compose with operation: "overlay". It places one still over one video and returns one MP4. It is a still, not animated motion graphics.
HeyGen's demo does something broader and live. Its September 2026 release post says liveavatar-hyperframes-demo has a background session write a HyperFrames composition that fades in over the avatar video, taking 30 seconds to two minutes. Sume facts are from Timeline compose, read 2026-10-01.
What does Sume compose overlay actually do?
Compose takes one still and one video, both already on this workspace's media.sume.com (import first with POST /v1/media-imports). Output length always comes from the video layer, and the still is held for the whole clip. The field that picks stack or overlay is operation, not mode.
How do I place the overlay?
| Key | Meaning for overlay |
|---|---|
position | top, center or bottom |
width_ratio | 0.05 to 1 of frame width (default 0.9); the plate keeps its aspect |
margin_ratio | 0 to 0.45 of height (default 0.05) |
video_fit | The only fit key allowed with overlay |
curl -X POST https://api.sume.com/v1/timeline-1.0/compose \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: overlay-001" \
-d '{"operation":"overlay",
"image":{"url":"https://media.sume.com/artifacts/artf_demo/plate.png"},
"video":{"url":"https://media.sume.com/artifacts/artf_demo/avatar.mp4"},
"layout":{"position":"bottom","width_ratio":0.8}}'What does it cost and how long does it take?
The docs give a public rate of $0.02 flat per job, to be confirmed live in GET /v1/catalog. Sync mode waits at most 30 seconds (wait_timeout_seconds is clamped to 0..30); otherwise poll the job.
What are the limits?
Current Avatar Video execution supports one resolved avatar per final video. Mixing stack keys into an overlay request is a 400. For a plate that states a label, see an AI-label compose example.
Sources
Related posts
More in Use cases
- HeyGen avatar new outfit with reference images vs Sume
HeyGen prompt avatars take avatar_id plus up to three reference_images for a new outfit. Sume's photo input takes one image_url per avatar.
- Add a hook title to the first seconds of a video with one cue
Send one authored cue with start 0 and end 3 to POST /v1/video-captions and Sume burns that hook text into the clip, with no speech-to-text step.
- Hue shift a video's colors by API for a Candy Pink look
Shift a whole video's hue and saturation with the hue, vibrance and colorize filters on Sume's video filter. One encode job, one new MP4, no preset library.
- Ideogram ad resizer API safe zones vs Sume aspect_ratio 4:5
Ideogram's ad resizer returns an exact WIDTHxHEIGHT inside a platform safe zone. Sume's image route picks an aspect_ratio and resolution tier instead.
Written by Sume