QA every episode's cliffhanger line with a video inspect transcript

Before you post episode 5, transcribe the render with POST /v1/video-inspect and check the last sentence. $0.01 per audio minute, plus a short Python gate.

5 min readSume
All posts

You can check that each rendered episode really ends on its cliffhanger line by transcribing the finished MP4 with POST /v1/video-inspect (transcribe: true) and comparing the last sentence of the transcript to the line in your script. The transcript rate is $0.01 per audio minute (read 2026-10-03), so an 8-episode season of 90-second cuts is 12 audio minutes, or $0.12 for the speech-to-text part.

The reason to automate it is the format. Metricool reports that TikTok and the agency Amplify launched "The Next Episode" to help fund creator-led series rather than one-off posts (reported, read 2026-10-03). A series ends every episode on a hook, and a model-generated clip is a stage where that hook can go missing: the line is cut by a trim, spoken in the wrong order, or replaced by something plausible but wrong.

What goes wrong at the end of an episode

Three failures show up at the tail of a cut. The last clip is shorter than the spoken line, so the sentence is clipped. The audio spine and the picture drift, so the line lands over the next scene. Or a regenerated clip carries different dialogue than the script, because the video model improvised. None of these is visible in a file listing, and all of them are cheap to catch with a transcript.

Video inspect reads one clip already hosted on media.sume.com, up to 1,800 seconds long, and with transcribe: true returns transcript.text, a words[] array, and optionally sentence segments[] when you ask for segmentation.mode: "sentence". It never re-encodes the clip and never makes a new MP4.

The request, field by field

Set frames to false so you get a probe and a transcript without stills. The docs list a probe-only inspect as enough to check probe.has_audio first, and a silent clip with transcribe: true fails with inspect_source_has_no_audio, which in a series gate is itself a useful signal.

video_inspect fields for a cliffhanger gate (read 2026-10-03)
FieldValueEffect
video_urlA media.sume.com artifactOff-host URLs are refused; import first
framesfalseProbe and transcript only, no stills
transcribetrueRuns Sume STT 1.0 on the audio
language_codeenA hint; omit to auto-detect
duration_seconds90Reserve hint, max 600; omit and it reserves one minute
segmentation{"mode": "sentence"}Adds sentence segments next to words

A gate you can run in CI

The script inspects an episode, takes the last sentence of the transcript, and fails unless it contains the hook phrase from your script. It compares plain text on purpose: STT may spell a name differently, so give it a short, distinctive phrase rather than the whole line.

import os, re, sys, time, requests

API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}

def transcript(video_url, key):
    body = {"video_url": video_url, "frames": False, "transcribe": True,
            "language_code": "en", "duration_seconds": 90}
    r = requests.post(f"{API}/v1/video-inspect",
                      headers={**H, "Idempotency-Key": key}, json=body)
    r.raise_for_status()
    rid = r.json()["request_id"]
    for _ in range(40):
        d = requests.get(f"{API}/v1/video-inspect/{rid}", headers=H).json()
        vi = d.get("video_inspect", d)
        if vi.get("transcript"):
            return vi["transcript"]["text"]
        time.sleep(3)
    raise TimeoutError(rid)

def last_sentence(text):
    parts = [p for p in re.split(r"(?<=[.!?])\s+", text.strip()) if p]
    return parts[-1] if parts else ""

ep, hook = sys.argv[1], sys.argv[2].lower()
tail = last_sentence(transcript(ep, "gate-" + ep[-12:]))
print("last sentence:", tail)
sys.exit(0 if hook in tail.lower() else 1)

What to do with the result

Run the gate on every episode after the final render and before the upload step, and store the transcript with the episode record. When it fails, look at the tail before you re-roll the whole episode: a clipped final line may be fixed by a longer last video slot in Timeline, and a drifted line by re-checking the audio spine start times, not by paying for new footage.

Budget the check separately from the render. The transcript part is $0.01 per audio minute. Video inspect also reserves its Modal compute ceiling and captures only the compute it used, so a probe-only inspect with frames: false is the light version; read the live figure in GET /v1/catalog before you scale it to a hundred episodes.

A worked example: episode 5 ends on the line "Nobody told her the door was open." Put nobody told her in the hook list. If the render ends on "Nobody told her the door was", the phrase still matches, so also compare the final word count to the script and flag any episode whose tail is more than a few words short. Two checks of different strictness catch both a missing hook and a clipped one.

Keep the hook phrases in the same file as the script, one per episode, so the gate and the writers share a source of truth. If the series changes a line late, the gate fails on purpose and tells you the rendered file is stale.

What Sume does and does not do

Sume returns a transcript of what is audible in the file and the probe facts of the clip. It does not judge whether a line is a good cliffhanger, and it does not return typed scenes from this route. For any semantic question about picture content, use the stills and look at them yourself.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume