Compare an avatar video's transcript_text to the approved script

Sume returns metadata.transcript_text on a finished avatar video once metadata is ready. Diff it against the approved script before publishing a support clip.

4 min readSume
All posts

Yes: a finished Sume avatar video carries metadata.transcript_text, a transcript of the rendered speech, so you can diff what the avatar said against the script a reviewer approved. The field fills in after the video completes, so read metadata.status first and only compare when it is ready.

This matters most for support and onboarding clips, where one dropped negation ("do not reset your password here") changes the meaning. A script review catches intent; a transcript diff catches what actually reached the file.

Where the transcript lives

Read the resource with GET /v1/avatar-videos/{id}. The response is data.avatar_video plus data.job. Inside avatar_video, metadata holds the enrichment: status, scenes, transcript_text, tags, summary, error, and timestamps. The OpenAPI describes this metadata as generated asynchronously after avatar-video completion, and detail readback returns the full shape by default.

Fields used for a transcript check (Sume OpenAPI, read 2026-10-05)
FieldTypeUse
avatar_video.resource_statusprocessing, ready, failed, canceled, archivedThe video itself is usable when ready
avatar_video.metadata.statusqueued, processing, ready, failedTranscript is trustworthy only when ready
avatar_video.metadata.transcript_textstring or nullWhat the render says
avatar_video.metadata.summary.transcript_word_countintegerCheap length sanity check

A 30-line check

The script below normalizes both texts and reports a similarity ratio plus the words that differ. The 0.92 threshold is our own starting heuristic, not a Sume value; tune it on your own clips.

import difflib, json, os, re, sys, urllib.request

API = "https://api.sume.com"

def get(path):
    req = urllib.request.Request(
        API + path,
        headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
                 "User-Agent": "transcript-check/1.0"})
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)

def words(text):
    return re.sub(r"[^a-z0-9' ]+", " ", text.lower()).split()

def check(video_id, approved):
    meta = get(f"/v1/avatar-videos/{video_id}")["data"]["avatar_video"]["metadata"]
    if not meta or meta["status"] != "ready" or not meta["transcript_text"]:
        return "metadata not ready, poll again"
    a, b = words(approved), words(meta["transcript_text"])
    ratio = difflib.SequenceMatcher(None, a, b).ratio()
    diff = [w for w in difflib.ndiff(a, b) if w[0] in "+-"]
    return ("PASS" if ratio >= 0.92 else "REVIEW", round(ratio, 3), diff[:20])

if __name__ == "__main__":
    print(check(sys.argv[1], open(sys.argv[2]).read()))

What to do with a mismatch

  • Small drift on numbers, product names, or acronyms: fix the script spelling (write the word the way it is spoken) and render again from a new first-frame preview.
  • A missing negation, price, or date: do not publish. Re-render; changed scripts need a new preview because structural fields cannot be edited on an approved one.
  • Transcript empty or metadata.status of failed: the video can still be fine. Watch it, or pull stills and audio with the video inspect route.

Limits of this check

transcript_text is a machine transcript of the finished file, so it can itself mis-hear a rare brand name. Treat a mismatch as a flag for a human listen, not as proof the avatar misspoke. It also says nothing about lip-sync quality or face drift; those need stills and a quick watch.

Keep the approved script, the avatar_video_id, and the ratio in your ticket or CMS record. If a clip is later questioned, you can show what was approved and what was rendered. Related reading: preview the first frame before paying for the render.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume