Put a logo or lower third over a talking avatar video with compose
Timeline compose overlay places a Sume-hosted still over an avatar clip, at the top, center or bottom. $0.02 flat per job. The layout keys and a request.
Use POST /v1/timeline-1.0/compose with operation 'overlay'. It puts one Sume-hosted still, such as a logo plate or a lower-third graphic, over a video, and you choose position top, center or bottom. It costs $0.02 flat per job.
What overlay does
Compose takes one still and one video and returns one MP4 with both on screen. In overlay the still sits on top of the video for the whole clip. The output length comes from the video layer, and the still can never make the clip longer.
Both URLs must already be on media.sume.com for your workspace, so import files first with POST /v1/media-imports. An Idempotency-Key header is required.
Layout keys for overlay
Mixing stack keys with overlay keys returns 400 (compose_overlay_takes_no_stack_layout), so send only these.
| Key | Allowed values | Default |
|---|---|---|
| position | top, center or bottom | Not stated on the page |
| width_ratio | 0.05 to 1 of the frame width | 0.9 |
| margin_ratio | 0 to 0.45 of the frame height | 0.05 |
| video_fit | cover, contain or stretch | Not stated on the page |
A lower third
A lower third wants a wide, short plate near the bottom edge. Make the PNG with transparency, import it, and then send this.
import json, os, urllib.request
body = {
"operation": "overlay",
"image": {"url": os.environ["PLATE_URL"]},
"video": {"url": os.environ["CLIP_URL"]},
"layout": {"position": "bottom", "width_ratio": 0.8, "margin_ratio": 0.06},
"mode": "sync",
}
req = urllib.request.Request(
"https://api.sume.com/v1/timeline-1.0/compose",
data=json.dumps(body).encode(),
method="POST",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json",
"Idempotency-Key": "lower-third-001",
},
)
with urllib.request.urlopen(req, timeout=40) as r:
print(r.status)Things to know
- A mute video returns a warning, compose_video_has_no_audio, not a failure.
- A video.duration past the end of the file is clamped with a warning.
- Default output is 1080 by 1920; set output.width and output.height to match your timeline so the shot is not rescaled twice.
- The overlay is a still, so it does not animate.
Cost
Branding ten avatar clips costs 10 x $0.02 = $0.20 in compose jobs. If a logo changes, re-compose from the clean clip; you do not need a new avatar render, which at plus quality would be $0.245 per second.
What to do
Keep the clean avatar clip and the plate as separate assets. That way a rebrand is a $0.02 job rather than a new render.
Sources
Related posts
More in Media tools
- Use a MAI-Voice-2.1 clip as avatar audio? Sume needs its own file
A talking still on Sume takes Sume-hosted audio under 10 MB. A MAI-Voice-2.1 clip sits elsewhere, so make the line with Sume TTS instead. Sizes and limits.
- Match voice emotion and music mood: one mood word, two Sume fields
Set generation_config.emotion on TTS and the emotion axis in the Music prompt from the same mood word, so a short's voice and bed do not argue.
- Microsoft's Content Provenance Detection: what you can check on a file
Foundry has a detection website and API for provenance. What the page says it checks, its limits, and how to use it on an AI clip or image from any generator.
- Microsoft: C2PA may not survive a crop, transcode or compression
Microsoft's provenance page lists the edits that can drop credentials and watermarks. Which of them a Sume trim, filter or caption job could be, and a check.
Written by Sume