Audit avatar videos for a missing CTA with cta_present
Sume's avatar-video metadata tags each scene with a segment_type and cta_present flag, so a script can list library clips that never ask viewers to act.
Fetch each ready avatar video and read metadata.summary.cta_present plus the per-scene cta_present and segment_type. Clips where no scene is a cta and the summary says false are your list of videos that explain a step but never tell the viewer what to do next. Sume computes these signals after the render, so you need no manual tagging.
What the metadata gives you
Per the Sume OpenAPI, each entry in metadata.scenes has timing (start_time_seconds, end_time_seconds), a segment_type, the scene transcript, on_screen_text, product_mentions, brand_mentions, a boolean-or-null cta_present, and a confidence. The summary rolls up scene_count, duration_seconds, cta_present, on_screen_text_present, and product_signals.
| Field | Where | Notes |
|---|---|---|
| segment_type | metadata.scenes[] | hook, problem, agitate, product_demo, product_reveal, proof, result, offer, cta, transition, other |
| cta_present | metadata.scenes[] | true, false, or null when unknown |
| cta_present | metadata.summary | Rolled up for the whole video |
| on_screen_text_present | metadata.summary | Useful for a caption and accessibility pass |
| confidence | metadata.scenes[] | low, medium, high, or a number |
Audit script
List ready videos, fetch each detail, and print those without a CTA. The list returns summaries only, so the detail read is required for metadata. Use the id from the list as the avatar-video id; the schema notes that id can be the linked job id until dedicated resource storage is ready, so treat a 404 as a skip.
import json, os, urllib.error, urllib.request
API = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
"User-Agent": "cta-audit/1.0"}
def get(path):
req = urllib.request.Request(API + path, headers=HEAD)
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["data"]
def main():
listing = get("/v1/avatar-videos?status=ready&limit=50")
for item in listing["avatar_videos"]:
try:
video = get("/v1/avatar-videos/" + item["id"])["avatar_video"]
except urllib.error.HTTPError:
continue
meta = video.get("metadata")
if not meta or meta["status"] != "ready":
continue
summary = meta.get("summary") or {}
kinds = [s["segment_type"] for s in meta["scenes"]]
if not summary.get("cta_present") and "cta" not in kinds:
print(video["id"], "no CTA", kinds)
main()Reading the result
- A support-answer clip without a CTA may be fine if the next action sits on the page around it. A sales or onboarding clip without one usually is not.
- If
cta_presentisnullon a scene, the metadata is unsure. Do not count it as a miss; check the transcript. - To fix a clip, add a closing
video_inputsscene with the action, then render again from a new preview. The multi-scene format is described in Generate avatar video.
Limits
These labels are generated signals, not a compliance record. A scene labelled cta is a model's reading of the transcript, so spot-check before you rewrite a whole library. The audit also only sees videos in your own workspace and only those whose metadata.status is ready. Videos rendered before you started relying on metadata may simply return null, so count them separately instead of treating them as clean.
Sources
Related posts
More in Sume Avatar 1.0
- Avatar captions failed with script_alignment_mismatch: what to change
A pinned captions.script_text that does not align with the speech soft-fails the captions; the clean video stays. Match the spoken words, then re-caption.
- Avatar Face Swap (Beta) for ad variants: 4 to 15 second source clips
Sume's Avatar Face Swap Beta puts a ready avatar face on a public source video. What it requires, what it does not take, and how to poll the job.
- Avatar inline captions vs standalone video captions: what is billed
Inline captions on a Sume avatar video create no separate caption job. Standalone captions are a billed job on a public URL. When each one is right.
- Avatar video aspect ratios: five choices, 720p only, 9:16 default
Sume Avatar 1.0 accepts 1:1, 3:4, 9:16, 4:3 and 16:9, defaults to 9:16, and renders 720p only. What each choice means for a vertical or landscape ad.
Written by Sume