Screen 30 holiday creator clips for the promo code with a transcript
Check creator-made holiday clips before they run: probe for audio and length, transcribe at $0.01 per audio minute, and search the words for the promo code.

To check that creator clips actually say your promo code, run POST /v1/video-inspect on each one with transcribe: true, then search transcript.text for the code. The transcript is Sume STT 1.0 at $0.01 per audio minute, so 30 clips of about a minute are about $0.30 for speech-to-text, plus each inspect's own compute, which is billed at its container seconds. Probing without a transcript costs no per-minute STT fee.
Q4 creator content arrives in volume and in every shape: a missing audio track, a clip that is too long, a code that was spoken wrongly. A programmatic check beats opening thirty files.
The result of a screening pass is a short list, not an edit. You still decide which clips run, and a person should watch the ones that pass before they are used in an ad.
What the call returns
Inspect reads one media.sume.com clip (import it with POST /v1/media-imports first) and returns a probe plus stills. With frames: false you get only the probe, which includes has_audio. With transcribe: true it also returns transcript with text, words[] and optional sentence segments[].
Do not send transcribe for a silent clip: the job fails with inspect_source_has_no_audio, so probe first. STT-only fields such as language_code need transcribe: true, or you get video_inspect_transcribe_required. The source limit is 1800 seconds and each call returns at most 24 stills.
Idempotency matters at this volume. Use a stable key per clip, such as the creator id plus the clip filename, so a retried batch does not start a second inspect for a clip that was already submitted.
| Check | Where it comes from | Cost |
|---|---|---|
| Has an audio track | probe.has_audio, frames false | Inspect compute only |
| Length within your limit | probe duration | Inspect compute only |
| Promo code spoken | transcript.text | $0.01 per audio minute, plus compute |
| Code on screen | stills at times you pick | Inspect compute only |
Matching the code
Set segmentation.mode: "sentence" if you want sentence segments[] with no gaps, in the shape of caption lines, and language_code to hint the language. Without duration_seconds Sume reserves one minute for the transcript, and the maximum hint is 600 seconds, so pass the real length for long clips.
Spoken codes are the weak point. Speech-to-text will mishear "BF25" as words, so normalise both sides before you compare: lower-case, strip spaces and punctuation, and also accept the code spelled out as letters. A fail is a prompt for a human to look, not proof the clip is wrong.
On-screen codes are a different check. Ask for stills at the seconds where the code should appear with frames: {at: [..]} (1 to 24 values) and read them yourself, or compare them against a reference, since a transcript cannot see text on screen.
A second pass can be cheaper than it sounds. Most failures show up in the probe alone, so run the probe on all 30, drop the clips that have no audio or are too long, and spend the transcript only on the ones that remain. That keeps the per-minute fee on clips that could actually run.
Probe first, then transcribe
Keep the screening script in your own repository and write its results to a spreadsheet, one row per clip: clip id, has_audio, duration, whether the code was found. A creator who is told exactly what failed fixes it in one round.
import os
import uuid
import requests
API = "https://api.sume.com/v1/video-inspect"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
CODE = "HOLIDAY25"
def inspect(url, **extra):
r = requests.post(API, headers={**H, "Idempotency-Key": f"scr-{uuid.uuid4()}"},
json={"video_url": url, **extra}, timeout=90)
return r.status_code, r.json()
url = os.environ["CREATOR_URL"]
code, probe = inspect(url, frames=False)
print(code, probe.get("request_id"))
# When the probe says has_audio, repeat the call with transcribe=True,
# then compare CODE with transcript.text after lower-casing both.
Mode and limits
The default mode is sync, which waits up to 30 seconds for a 200, and otherwise returns 202 and a job to poll with GET /v1/jobs/:id/status. For thirty clips, submit them with a small concurrency and poll, rather than waiting on each.
A transcript is a private working file. It is not a rights clearance, and it does not tell you whether the creator may use the music in the clip. Check usage terms with the creator before anything goes into a paid placement, and add any disclosure the platform requires for sponsored posts.
Sources
Related posts
More in Use cases
- Holiday video drop at 9am: submit by when? The Sume 90-minute deadline
A Format run is force-finalized as failed 90 minutes after creation, so a 9am drop needs the batch submitted well before 7:30, with headroom for retries.
- Hotel room tour in one take: 30 seconds of Seedance 2.5 from photos
Turn room photos into a 30-second one-take hotel tour on Seedance 2.5: prompt route, first-frame photo, what it invents, and the Sume price at 720p and 1080p.
- Houseplant shop montage with no voice: a silent spine and a Lyria bed
Make a music-only product montage: Timeline audio.mode silence, a Lyria 3.5 bed as soundtrack and eight plant clips, for $0.225 on Sume.
- How fast is a reference Short cut? Shot length from reference-ingest
Average shot length is clip duration divided by the shots[] count in a Sume reference-ingest manifest. It is a detector estimate, not an editor's cut list.
Written by Sume