Avatar video is ready but transcript_text is null: poll metadata again
A Sume avatar video can be resource_status ready while metadata.status is still processing. The file is usable now; the transcript and tags arrive later.
If GET /v1/avatar-videos/{id} shows resource_status: "ready" but metadata.transcript_text is null, nothing is wrong: the video is finished and the enrichment is a separate, later step. Download or publish the MP4 from video_url as soon as it is ready, and poll the same resource again for the transcript, tags, and scene breakdown.
Two clocks on one resource
The Sume OpenAPI describes metadata as rich public-safe video metadata generated asynchronously after avatar-video completion. So a finished clip has two independent states: the render (resource_status and job_status) and the enrichment (metadata.status). Code that waits for both on one condition will block longer than it needs to, and code that reads the transcript the moment the render completes will see null.
| Field | Values | Meaning |
|---|---|---|
| resource_status | processing, ready, failed, canceled, archived | Completed jobs map to ready |
| job_status | queued, processing, completed, failed, canceled | Underlying job state, or null with no linked job |
| metadata.status | queued, processing, ready, failed | Enrichment of transcript, tags, scenes, summary |
| captions.status (if requested) | skipped, pending, ready, failed | Caption burn-in is its own stage and can soft-fail |
What to gate on
- Publishing the clip: gate on
resource_status == "ready"and a non-nullvideo_url. If you asked for captions, also look atcaptions.statusto know whethervideo_urlis the captioned or the clean file. - Search, tagging, or a transcript diff: gate on
metadata.status == "ready". - Alerting: a
metadata.statusoffailedcarries anerrorobject, but it does not turn a ready video into a failed one.
A bounded wait for metadata
Poll with backoff and a ceiling so a stuck enrichment never hangs a pipeline. The docs tell you to use exponential backoff for job status; the same habit applies here. This sketch gives up after a fixed number of tries and returns what it has.
import json, os, time, urllib.request
API = "https://api.sume.com"
def get_video(video_id):
req = urllib.request.Request(
f"{API}/v1/avatar-videos/{video_id}",
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"User-Agent": "metadata-wait/1.0"})
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]["avatar_video"]
def wait_for_metadata(video_id, tries=8, first_delay=5):
delay = first_delay
video = get_video(video_id)
for _ in range(tries):
meta = video.get("metadata")
if video["resource_status"] in ("failed", "canceled"):
return video
if meta and meta["status"] in ("ready", "failed"):
return video
time.sleep(delay)
delay = min(delay * 2, 60)
video = get_video(video_id)
return videoPitfalls
Do not resubmit the avatar request because the transcript is missing. That would bill a second render for a problem that is only a timing gap. The jobs documentation is explicit that you must not submit a new paid job for the same intent; a retry of the submit itself should reuse the same Idempotency-Key.
The metadata object can also be null on a resource that has no enrichment yet, so test for it before indexing into it, as the code above does. Sume does not promise a specific enrichment time in the docs, so keep the ceiling configurable and log the avatar_video_id of anything that exceeds it.
Sources
Related posts
More in Sume Avatar 1.0
- Compare an avatar video's transcript_text to the approved script
Sume returns metadata.transcript_text on a finished avatar video once metadata is ready. Diff it against the approved script before publishing a support clip.
- Bystander faces in a scene photo: check likeness before image_url
A photo scene can carry a stranger's face into the render. Check every face before you send image_url; the docs do not describe a screening step.
- Create a holiday host avatar once: prompt, profile or photo, $0.95
Create one reusable Sume avatar from a prompt, a profile or a photo for $0.95, then reuse its handle in every holiday clip. Inputs and a photo request.
- Does an AI avatar presenter make a Short original?
An avatar is a delivery method; YouTube's pages ask for original substance. How to give an Avatar 1.0 Short your own angle within its 60-second job limit.
Written by Sume