QA every episode's cliffhanger line with a video inspect transcript
Before you post episode 5, transcribe the render with POST /v1/video-inspect and check the last sentence. $0.01 per audio minute, plus a short Python gate.

You can check that each rendered episode really ends on its cliffhanger line by transcribing the finished MP4 with POST /v1/video-inspect (transcribe: true) and comparing the last sentence of the transcript to the line in your script. The transcript rate is $0.01 per audio minute (read 2026-10-03), so an 8-episode season of 90-second cuts is 12 audio minutes, or $0.12 for the speech-to-text part.
The reason to automate it is the format. Metricool reports that TikTok and the agency Amplify launched "The Next Episode" to help fund creator-led series rather than one-off posts (reported, read 2026-10-03). A series ends every episode on a hook, and a model-generated clip is a stage where that hook can go missing: the line is cut by a trim, spoken in the wrong order, or replaced by something plausible but wrong.
What goes wrong at the end of an episode
Three failures show up at the tail of a cut. The last clip is shorter than the spoken line, so the sentence is clipped. The audio spine and the picture drift, so the line lands over the next scene. Or a regenerated clip carries different dialogue than the script, because the video model improvised. None of these is visible in a file listing, and all of them are cheap to catch with a transcript.
Video inspect reads one clip already hosted on media.sume.com, up to 1,800 seconds long, and with transcribe: true returns transcript.text, a words[] array, and optionally sentence segments[] when you ask for segmentation.mode: "sentence". It never re-encodes the clip and never makes a new MP4.
The request, field by field
Set frames to false so you get a probe and a transcript without stills. The docs list a probe-only inspect as enough to check probe.has_audio first, and a silent clip with transcribe: true fails with inspect_source_has_no_audio, which in a series gate is itself a useful signal.
| Field | Value | Effect |
|---|---|---|
| video_url | A media.sume.com artifact | Off-host URLs are refused; import first |
| frames | false | Probe and transcript only, no stills |
| transcribe | true | Runs Sume STT 1.0 on the audio |
| language_code | en | A hint; omit to auto-detect |
| duration_seconds | 90 | Reserve hint, max 600; omit and it reserves one minute |
| segmentation | {"mode": "sentence"} | Adds sentence segments next to words |
A gate you can run in CI
The script inspects an episode, takes the last sentence of the transcript, and fails unless it contains the hook phrase from your script. It compares plain text on purpose: STT may spell a name differently, so give it a short, distinctive phrase rather than the whole line.
import os, re, sys, time, requests
API = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
def transcript(video_url, key):
body = {"video_url": video_url, "frames": False, "transcribe": True,
"language_code": "en", "duration_seconds": 90}
r = requests.post(f"{API}/v1/video-inspect",
headers={**H, "Idempotency-Key": key}, json=body)
r.raise_for_status()
rid = r.json()["request_id"]
for _ in range(40):
d = requests.get(f"{API}/v1/video-inspect/{rid}", headers=H).json()
vi = d.get("video_inspect", d)
if vi.get("transcript"):
return vi["transcript"]["text"]
time.sleep(3)
raise TimeoutError(rid)
def last_sentence(text):
parts = [p for p in re.split(r"(?<=[.!?])\s+", text.strip()) if p]
return parts[-1] if parts else ""
ep, hook = sys.argv[1], sys.argv[2].lower()
tail = last_sentence(transcript(ep, "gate-" + ep[-12:]))
print("last sentence:", tail)
sys.exit(0 if hook in tail.lower() else 1)What to do with the result
Run the gate on every episode after the final render and before the upload step, and store the transcript with the episode record. When it fails, look at the tail before you re-roll the whole episode: a clipped final line may be fixed by a longer last video slot in Timeline, and a drifted line by re-checking the audio spine start times, not by paying for new footage.
Budget the check separately from the render. The transcript part is $0.01 per audio minute. Video inspect also reserves its Modal compute ceiling and captures only the compute it used, so a probe-only inspect with frames: false is the light version; read the live figure in GET /v1/catalog before you scale it to a hundred episodes.
A worked example: episode 5 ends on the line "Nobody told her the door was open." Put nobody told her in the hook list. If the render ends on "Nobody told her the door was", the phrase still matches, so also compare the final word count to the script and flag any episode whose tail is more than a few words short. Two checks of different strictness catch both a missing hook and a clipped one.
Keep the hook phrases in the same file as the script, one per episode, so the gate and the writers share a source of truth. If the series changes a line late, the gate fails on purpose and tells you the rendered file is stale.
What Sume does and does not do
Sume returns a transcript of what is audible in the file and the probe facts of the clip. It does not judge whether a line is a good cliffhanger, and it does not return typed scenes from this route. For any semantic question about picture content, use the stills and look at them yourself.
Sources
Related posts
More in Use cases
- Qwen Image Max backdrop plates for product composites
Generate empty lifestyle backdrops with Qwen Image Max, then place the real product photo: vendor facts, price, and Sume's text-only catalog entry.
- Rakuten Ichiba image rule: text on 20% or less of the photo
Rakuten Ichiba's image guideline: text on 20% or less, white or photo background, no frames, no GIF, square 700x700. Make the base text-free on Sume.
- How to make a reaction Short that counts as original value (Oct 2026)
YouTube now favors Shorts with your own commentary over reposts. What a reaction Short needs, and which steps Sume can cut, caption and assemble for you.
- Real estate brokerage: one listing video per property in a bulk run
Queue up to 100 listing videos in one Format bulk run: per-item inputs, a spend cap on each item, and reading counts.failed instead of trusting completed.
Written by Sume