Verify a spoken ad disclosure is in the final mix with STT word times

Detach the final ad's audio, transcribe it with Sume STT, and confirm the required line is spoken and when it starts. Python check, not legal advice.

4 min readSume
All posts

Detach the audio from the finished ad as 16 kHz mono wav, transcribe it with Sume STT, and search the words array for your approved disclosure wording. The first matching word's start time says when the line is heard. If no match comes back, listen to the mix before it ships.

This checks that your own audio contains a line. It does not decide whether a platform, client or law requires that line, and it is not legal advice.

Why check the final mix and not the script?

A line can be in the script and still be lost: a music bed that is too loud, a trim that cuts the end, or a retake that changed a word. The only file a viewer hears is the final render, so test that file.

How do I get the audio and the words?

Audio detach takes one media.sume.com video and returns a new audio artifact; the default is wav, and mono at 16 kHz is the shape the docs call the STT shape. It costs $0.01 per job. Then POST /v1/stt-1.0/transcribe with audio_url and, for an English ad, language_code set to en. Add duration_seconds if you know it. A source with no audio track fails with detach_source_has_no_audio.

STT is a model and can mishear. If a phrase is missing, treat that as a prompt to listen, not as proof.

What the check tests (Sume docs and OpenAPI, read 2026-10-02).
QuestionWhere to look
Is the line spoken at all?Normalized words[].word contains your wording in order
When does it start?start of the first matching word
Is it all there?end of the last matching word
Is it early enough?Compare the start with your own limit

What does the matcher look like?

The sketch lowercases, strips punctuation (so a hyphenated token splits into two words) and looks for the wording as a run. It runs as written.

import re

required = "this ad uses an ai generated voice"   # your approved wording
words = [("This", 1.2), ("ad", 1.4), ("uses", 1.6), ("an", 1.8), ("AI-generated", 1.9),
         ("voice", 2.5)]
tokens = [(re.sub(r"[^a-z0-9]+", " ", w.lower()).strip(), t) for w, t in words]
flat = [(p, t) for w, t in tokens for p in w.split()]
need = required.split()
for i in range(len(flat) - len(need) + 1):
    if [p for p, _ in flat[i:i + len(need)]] == need:
        print("found at", flat[i][1], "s; ends", flat[i + len(need) - 1][1], "s")
        break
else:
    print("NOT FOUND: listen to the mix")

What else belongs in the record?

Save the STT job id, the matched start time and the wording you tested with the ad's job ids. If the wording is on-screen rather than spoken, check that separately with video frames. Where a rule comes from a platform page, read it on that platform's own page; this check only proves your file matches your wording.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume