Audit avatar videos for a missing CTA with cta_present

Sume's avatar-video metadata tags each scene with a segment_type and cta_present flag, so a script can list library clips that never ask viewers to act.

4 min readSume
All posts

Fetch each ready avatar video and read metadata.summary.cta_present plus the per-scene cta_present and segment_type. Clips where no scene is a cta and the summary says false are your list of videos that explain a step but never tell the viewer what to do next. Sume computes these signals after the render, so you need no manual tagging.

What the metadata gives you

Per the Sume OpenAPI, each entry in metadata.scenes has timing (start_time_seconds, end_time_seconds), a segment_type, the scene transcript, on_screen_text, product_mentions, brand_mentions, a boolean-or-null cta_present, and a confidence. The summary rolls up scene_count, duration_seconds, cta_present, on_screen_text_present, and product_signals.

Scene and summary fields for a CTA audit (Sume OpenAPI, read 2026-10-05)
FieldWhereNotes
segment_typemetadata.scenes[]hook, problem, agitate, product_demo, product_reveal, proof, result, offer, cta, transition, other
cta_presentmetadata.scenes[]true, false, or null when unknown
cta_presentmetadata.summaryRolled up for the whole video
on_screen_text_presentmetadata.summaryUseful for a caption and accessibility pass
confidencemetadata.scenes[]low, medium, high, or a number

Audit script

List ready videos, fetch each detail, and print those without a CTA. The list returns summaries only, so the detail read is required for metadata. Use the id from the list as the avatar-video id; the schema notes that id can be the linked job id until dedicated resource storage is ready, so treat a 404 as a skip.

import json, os, urllib.error, urllib.request

API = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"],
        "User-Agent": "cta-audit/1.0"}

def get(path):
    req = urllib.request.Request(API + path, headers=HEAD)
    with urllib.request.urlopen(req, timeout=30) as r:
        return json.load(r)["data"]

def main():
    listing = get("/v1/avatar-videos?status=ready&limit=50")
    for item in listing["avatar_videos"]:
        try:
            video = get("/v1/avatar-videos/" + item["id"])["avatar_video"]
        except urllib.error.HTTPError:
            continue
        meta = video.get("metadata")
        if not meta or meta["status"] != "ready":
            continue
        summary = meta.get("summary") or {}
        kinds = [s["segment_type"] for s in meta["scenes"]]
        if not summary.get("cta_present") and "cta" not in kinds:
            print(video["id"], "no CTA", kinds)

main()

Reading the result

  • A support-answer clip without a CTA may be fine if the next action sits on the page around it. A sales or onboarding clip without one usually is not.
  • If cta_present is null on a scene, the metadata is unsure. Do not count it as a miss; check the transcript.
  • To fix a clip, add a closing video_inputs scene with the action, then render again from a new preview. The multi-scene format is described in Generate avatar video.

Limits

These labels are generated signals, not a compliance record. A scene labelled cta is a model's reading of the transcript, so spot-check before you rewrite a whole library. The audit also only sees videos in your own workspace and only those whose metadata.status is ready. Videos rendered before you started relying on metadata may simply return null, so count them separately instead of treating them as clean.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume