Compare an avatar video's transcript_text to the approved script
Sume returns metadata.transcript_text on a finished avatar video once metadata is ready. Diff it against the approved script before publishing a support clip.
Yes: a finished Sume avatar video carries metadata.transcript_text, a transcript of the rendered speech, so you can diff what the avatar said against the script a reviewer approved. The field fills in after the video completes, so read metadata.status first and only compare when it is ready.
This matters most for support and onboarding clips, where one dropped negation ("do not reset your password here") changes the meaning. A script review catches intent; a transcript diff catches what actually reached the file.
Where the transcript lives
Read the resource with GET /v1/avatar-videos/{id}. The response is data.avatar_video plus data.job. Inside avatar_video, metadata holds the enrichment: status, scenes, transcript_text, tags, summary, error, and timestamps. The OpenAPI describes this metadata as generated asynchronously after avatar-video completion, and detail readback returns the full shape by default.
| Field | Type | Use |
|---|---|---|
| avatar_video.resource_status | processing, ready, failed, canceled, archived | The video itself is usable when ready |
| avatar_video.metadata.status | queued, processing, ready, failed | Transcript is trustworthy only when ready |
| avatar_video.metadata.transcript_text | string or null | What the render says |
| avatar_video.metadata.summary.transcript_word_count | integer | Cheap length sanity check |
A 30-line check
The script below normalizes both texts and reports a similarity ratio plus the words that differ. The 0.92 threshold is our own starting heuristic, not a Sume value; tune it on your own clips.
import difflib, json, os, re, sys, urllib.request
API = "https://api.sume.com"
def get(path):
req = urllib.request.Request(
API + path,
headers={"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"User-Agent": "transcript-check/1.0"})
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)
def words(text):
return re.sub(r"[^a-z0-9' ]+", " ", text.lower()).split()
def check(video_id, approved):
meta = get(f"/v1/avatar-videos/{video_id}")["data"]["avatar_video"]["metadata"]
if not meta or meta["status"] != "ready" or not meta["transcript_text"]:
return "metadata not ready, poll again"
a, b = words(approved), words(meta["transcript_text"])
ratio = difflib.SequenceMatcher(None, a, b).ratio()
diff = [w for w in difflib.ndiff(a, b) if w[0] in "+-"]
return ("PASS" if ratio >= 0.92 else "REVIEW", round(ratio, 3), diff[:20])
if __name__ == "__main__":
print(check(sys.argv[1], open(sys.argv[2]).read()))What to do with a mismatch
- Small drift on numbers, product names, or acronyms: fix the script spelling (write the word the way it is spoken) and render again from a new first-frame preview.
- A missing negation, price, or date: do not publish. Re-render; changed scripts need a new preview because structural fields cannot be edited on an approved one.
- Transcript empty or
metadata.statusoffailed: the video can still be fine. Watch it, or pull stills and audio with the video inspect route.
Limits of this check
transcript_text is a machine transcript of the finished file, so it can itself mis-hear a rare brand name. Treat a mismatch as a flag for a human listen, not as proof the avatar misspoke. It also says nothing about lip-sync quality or face drift; those need stills and a quick watch.
Keep the approved script, the avatar_video_id, and the ratio in your ticket or CMS record. If a clip is later questioned, you can show what was approved and what was rendered. Related reading: preview the first frame before paying for the render.
Sources
Related posts
More in Sume Avatar 1.0
- Bystander faces in a scene photo: check likeness before image_url
A photo scene can carry a stranger's face into the render. Check every face before you send image_url; the docs do not describe a screening step.
- Create a holiday host avatar once: prompt, profile or photo, $0.95
Create one reusable Sume avatar from a prompt, a profile or a photo for $0.95, then reuse its handle in every holiday clip. Inputs and a photo request.
- Does an AI avatar presenter make a Short original?
An avatar is a delivery method; YouTube's pages ask for original substance. How to give an Avatar 1.0 Short your own angle within its 60-second job limit.
- Face swap beta: motion stage runs on Kling, source audio muxed back
Sume's Avatar Face Swap beta runs its motion stage on the Kling 3.0 motion control queue, strips the source audio, then muxes the original audio back in.
Written by Sume