Avatar package with captions and soundtrack: which video_url you get
In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.
With package, Sume runs post-production in a fixed order: clean talking-head, then captions on the clean video, then the soundtrack mixed under the captioned file. If captions or the soundtrack fail, the job does not fail. video_url points to the furthest successful stage: the captioned MP4 if captions are ready, otherwise the clean talking-head.
The order, from the OpenAPI
The create request accepts an optional package object with captions and soundtrack. The contract says caption timing is aligned on the clean video, never after music is mixed, so music cannot confuse caption timing. A soundtrack comes from either soundtrack.prompt (generated with Music 1.0) or soundtrack.audio_url (your own public HTTPS audio, mirrored and mixed with no Music call). Default linear volume is 0.15. Package captions are mutually exclusive with top-level captions when package.captions.enabled is set.
| Captions | Soundtrack and mux | video_url |
|---|---|---|
| ready | ready | Captioned video with music ducked under speech |
| ready | failed | Captioned video, no music |
| failed | ready | Clean video with music |
| failed | failed | Clean talking-head |
| not requested | ready | Clean video with music |
Read the stage fields, not just video_url
The resource returns package.captions, package.soundtrack, and package.mux, each with a status of skipped, pending, ready, or failed. A failed stage carries a typed public_reason, for example soundtrack_unavailable, music_generation_failed, missing_music_audio, or a caption reason such as script_alignment_mismatch. While a stage is pending, an audio_url you supplied is echoed back so you can confirm intent; a prompt soundtrack stays null until ready.
import json, os, urllib.request
def describe(video_id):
req = urllib.request.Request(
f"https://api.sume.com/v1/avatar-videos/{video_id}",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"User-Agent": "package-check/1.0"})
with urllib.request.urlopen(req, timeout=30) as r:
video = json.load(r)["data"]["avatar_video"]
pkg = video.get("package") or {}
for stage in ("captions", "soundtrack", "mux"):
info = pkg.get(stage) or {}
print(stage, info.get("status", "not requested"), info.get("public_reason", ""))
print("primary:", video["video_url"])Practical rules
- Branch on the stage statuses before you tell a customer the clip is branded. A clean file labelled as captioned is the most likely mistake.
- If captions failed and you need them, recaption the clean video with the standalone route after fixing the cause. If only the soundtrack failed, you can mix music in a timeline compose step instead of re-rendering the avatar.
- Captions require speakable script text and reject an estimated duration above 60 seconds, the same ceiling as the avatar window.
- The inline caption add-on is a fixed $0.20 per the OpenAPI and is part of the avatar-video usage estimate when enabled, so a soft-failed caption stage is worth checking on your usage line.
When not to use package
Skip the package if the clip goes into an editor or a timeline later. Burned-in captions and a music bed cannot be removed from the MP4, and music mixed too early makes any later mix harder. Render the clean avatar, then caption and score it where you control the final cut.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar video with a product image: the premium is 1-3 cents a second
A product_image on a Sume Avatar 1.0 video adds $0.010 (standard), $0.013 (plus) or $0.030 (max) per second: 30 to 90 cents on a 30-second ad.
- Avatar video is ready but transcript_text is null: poll metadata again
A Sume avatar video can be resource_status ready while metadata.status is still processing. The file is usable now; the transcript and tags arrive later.
- Compare an avatar video's transcript_text to the approved script
Sume returns metadata.transcript_text on a finished avatar video once metadata is ready. Diff it against the approved script before publishing a support clip.
- Bystander faces in a scene photo: check likeness before image_url
A photo scene can carry a stranger's face into the render. Check every face before you send image_url; the docs do not describe a screening step.
Written by Sume