Avatar video is ready but transcript_text is null: poll metadata again

A Sume avatar video can be resource_status ready while metadata.status is still processing. The file is usable now; the transcript and tags arrive later.

4 min readSume
All posts

If GET /v1/avatar-videos/{id} shows resource_status: "ready" but metadata.transcript_text is null, nothing is wrong: the video is finished and the enrichment is a separate, later step. Download or publish the MP4 from video_url as soon as it is ready, and poll the same resource again for the transcript, tags, and scene breakdown.

Two clocks on one resource

The Sume OpenAPI describes metadata as rich public-safe video metadata generated asynchronously after avatar-video completion. So a finished clip has two independent states: the render (resource_status and job_status) and the enrichment (metadata.status). Code that waits for both on one condition will block longer than it needs to, and code that reads the transcript the moment the render completes will see null.

State fields on an avatar video (Sume OpenAPI, read 2026-10-05)
FieldValuesMeaning
resource_statusprocessing, ready, failed, canceled, archivedCompleted jobs map to ready
job_statusqueued, processing, completed, failed, canceledUnderlying job state, or null with no linked job
metadata.statusqueued, processing, ready, failedEnrichment of transcript, tags, scenes, summary
captions.status (if requested)skipped, pending, ready, failedCaption burn-in is its own stage and can soft-fail

What to gate on

  • Publishing the clip: gate on resource_status == "ready" and a non-null video_url. If you asked for captions, also look at captions.status to know whether video_url is the captioned or the clean file.
  • Search, tagging, or a transcript diff: gate on metadata.status == "ready".
  • Alerting: a metadata.status of failed carries an error object, but it does not turn a ready video into a failed one.

A bounded wait for metadata

Poll with backoff and a ceiling so a stuck enrichment never hangs a pipeline. The docs tell you to use exponential backoff for job status; the same habit applies here. This sketch gives up after a fixed number of tries and returns what it has.

import json, os, time, urllib.request

API = "https://api.sume.com"

def get_video(video_id):
    req = urllib.request.Request(
        f"{API}/v1/avatar-videos/{video_id}",
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "User-Agent": "metadata-wait/1.0"})
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)["data"]["avatar_video"]

def wait_for_metadata(video_id, tries=8, first_delay=5):
    delay = first_delay
    video = get_video(video_id)
    for _ in range(tries):
        meta = video.get("metadata")
        if video["resource_status"] in ("failed", "canceled"):
            return video
        if meta and meta["status"] in ("ready", "failed"):
            return video
        time.sleep(delay)
        delay = min(delay * 2, 60)
        video = get_video(video_id)
    return video

Pitfalls

Do not resubmit the avatar request because the transcript is missing. That would bill a second render for a problem that is only a timing gap. The jobs documentation is explicit that you must not submit a new paid job for the same intent; a retry of the submit itself should reuse the same Idempotency-Key.

The metadata object can also be null on a resource that has no enrichment yet, so test for it before indexing into it, as the code above does. Sume does not promise a specific enrichment time in the docs, so keep the ceiling configurable and log the avatar_video_id of anything that exceeds it.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume