Check that a recorded ad says the required disclaimer: STT words

Transcribe an ad with Sume STT, match the disclaimer in words[], and fail the check if it is missing or starts in the last seconds. Python, 1 cent per minute.

4 min readSume
All posts

A finance, health or promotion ad often has a line that has to be spoken, such as "Terms and conditions apply" or "Results may vary". Editors cut it by accident, and a voiceover swap can drop it. A check at export time is cheaper than a takedown. Sume STT 1.0 gives you words[] with times, which is all the check needs.

What to test

Two things matter: the words are there in order, and the viewer can hear them. The first test is a phrase match over the words with type == "word". The second is a simple timing rule: the phrase must start before the final 1.5 seconds, so it does not get clipped at the end of the video. Sume does not judge whether the wording is legally sufficient. Your compliance owner decides that. The code only proves the words were spoken.

Code

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

import re
def norm(t):
    return re.sub(r"[^a-z0-9']", "", t.lower())

def has_phrase(words, phrase, end_total, tail=1.5):
    toks = [(norm(w["word"]), w) for w in words if w.get("type") == "word"]
    want = [norm(p) for p in phrase.split()]
    for i in range(len(toks) - len(want) + 1):
        if [t for t, _ in toks[i:i + len(want)]] == want:
            return toks[i][1]["start"] < end_total - tail
    return False

res = run("/stt-1.0/transcribe", {"audio_url": os.environ["AD_URL"], "duration_seconds": 60})
end = res["words"][-1]["end"]
print("OK" if has_phrase(res["words"], "terms and conditions apply", end) else "MISSING")

Cost and limits

STT 1.0 costs about $0.01 per minute of audio, so a 30-second ad is about half a cent of usage, and duration_seconds is a reservation hint from 1 to 600 that sizes the quote. A job takes at most 10 minutes of audio, and audio_url must be public HTTPS. For a batch of 200 ad variants of 30 seconds, that is 100 minutes, or about $1 before any rounding of small charges.

Make the match tolerant, not loose

Transcription can write "T and C's" or "terms & conditions". Keep a short list of accepted spellings, normalized the same way. Do not use fuzzy matching broad enough to accept a different sentence. If the check fails, hand the clip to a person rather than auto-fixing it. If the language is not English, pass language_code so the provider does not have to guess.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume