Avatar package with captions and soundtrack: which video_url you get

In a Sume avatar package, captions burn onto the clean video first, then music is mixed in. If a stage soft-fails, video_url is the furthest successful file.

4 min readSume
All posts

With package, Sume runs post-production in a fixed order: clean talking-head, then captions on the clean video, then the soundtrack mixed under the captioned file. If captions or the soundtrack fail, the job does not fail. video_url points to the furthest successful stage: the captioned MP4 if captions are ready, otherwise the clean talking-head.

The order, from the OpenAPI

The create request accepts an optional package object with captions and soundtrack. The contract says caption timing is aligned on the clean video, never after music is mixed, so music cannot confuse caption timing. A soundtrack comes from either soundtrack.prompt (generated with Music 1.0) or soundtrack.audio_url (your own public HTTPS audio, mirrored and mixed with no Music call). Default linear volume is 0.15. Package captions are mutually exclusive with top-level captions when package.captions.enabled is set.

What video_url points to after each outcome (Sume OpenAPI, read 2026-10-05)
CaptionsSoundtrack and muxvideo_url
readyreadyCaptioned video with music ducked under speech
readyfailedCaptioned video, no music
failedreadyClean video with music
failedfailedClean talking-head
not requestedreadyClean video with music

Read the stage fields, not just video_url

The resource returns package.captions, package.soundtrack, and package.mux, each with a status of skipped, pending, ready, or failed. A failed stage carries a typed public_reason, for example soundtrack_unavailable, music_generation_failed, missing_music_audio, or a caption reason such as script_alignment_mismatch. While a stage is pending, an audio_url you supplied is echoed back so you can confirm intent; a prompt soundtrack stays null until ready.

import json, os, urllib.request

def describe(video_id):
    req = urllib.request.Request(
        f"https://api.sume.com/v1/avatar-videos/{video_id}",
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "User-Agent": "package-check/1.0"})
    with urllib.request.urlopen(req, timeout=30) as r:
        video = json.load(r)["data"]["avatar_video"]
    pkg = video.get("package") or {}
    for stage in ("captions", "soundtrack", "mux"):
        info = pkg.get(stage) or {}
        print(stage, info.get("status", "not requested"), info.get("public_reason", ""))
    print("primary:", video["video_url"])

Practical rules

  • Branch on the stage statuses before you tell a customer the clip is branded. A clean file labelled as captioned is the most likely mistake.
  • If captions failed and you need them, recaption the clean video with the standalone route after fixing the cause. If only the soundtrack failed, you can mix music in a timeline compose step instead of re-rendering the avatar.
  • Captions require speakable script text and reject an estimated duration above 60 seconds, the same ceiling as the avatar window.
  • The inline caption add-on is a fixed $0.20 per the OpenAPI and is part of the avatar-video usage estimate when enabled, so a soft-failed caption stage is worth checking on your usage line.

When not to use package

Skip the package if the clip goes into an editor or a timeline later. Burned-in captions and a music bed cannot be removed from the MP4, and music mixed too early makes any later mix harder. Render the clean avatar, then caption and score it where you control the final cut.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume