Verify a spoken ad disclosure is in the final mix with STT word times
Detach the final ad's audio, transcribe it with Sume STT, and confirm the required line is spoken and when it starts. Python check, not legal advice.

Detach the audio from the finished ad as 16 kHz mono wav, transcribe it with Sume STT, and search the words array for your approved disclosure wording. The first matching word's start time says when the line is heard. If no match comes back, listen to the mix before it ships.
This checks that your own audio contains a line. It does not decide whether a platform, client or law requires that line, and it is not legal advice.
Why check the final mix and not the script?
A line can be in the script and still be lost: a music bed that is too loud, a trim that cuts the end, or a retake that changed a word. The only file a viewer hears is the final render, so test that file.
How do I get the audio and the words?
Audio detach takes one media.sume.com video and returns a new audio artifact; the default is wav, and mono at 16 kHz is the shape the docs call the STT shape. It costs $0.01 per job. Then POST /v1/stt-1.0/transcribe with audio_url and, for an English ad, language_code set to en. Add duration_seconds if you know it. A source with no audio track fails with detach_source_has_no_audio.
STT is a model and can mishear. If a phrase is missing, treat that as a prompt to listen, not as proof.
| Question | Where to look |
|---|---|
| Is the line spoken at all? | Normalized words[].word contains your wording in order |
| When does it start? | start of the first matching word |
| Is it all there? | end of the last matching word |
| Is it early enough? | Compare the start with your own limit |
What does the matcher look like?
The sketch lowercases, strips punctuation (so a hyphenated token splits into two words) and looks for the wording as a run. It runs as written.
import re
required = "this ad uses an ai generated voice" # your approved wording
words = [("This", 1.2), ("ad", 1.4), ("uses", 1.6), ("an", 1.8), ("AI-generated", 1.9),
("voice", 2.5)]
tokens = [(re.sub(r"[^a-z0-9]+", " ", w.lower()).strip(), t) for w, t in words]
flat = [(p, t) for w, t in tokens for p in w.split()]
need = required.split()
for i in range(len(flat) - len(need) + 1):
if [p for p, _ in flat[i:i + len(need)]] == need:
print("found at", flat[i][1], "s; ends", flat[i + len(need) - 1][1], "s")
break
else:
print("NOT FOUND: listen to the mix")
What else belongs in the record?
Save the STT job id, the matched start time and the wording you tested with the ad's job ids. If the wording is on-screen rather than spoken, check that separately with video frames. Where a rule comes from a platform page, read it on that platform's own page; this check only proves your file matches your wording.
Sources
Related posts
More in Use cases
- Vinted bans edited and stock photos: what AI can still do
Vinted's rules say photos must show the item as it is with no image editing, and ban stock and watermarked pictures. Keep AI out of listings; use it for promos.
- Vocabulary flashcard video: one still and one spoken word per card
Build a vocabulary flashcard video: one image and one TTS word per card, joined with Timeline audio concat, then re-based into render slots from segments[].
- Voice replication blocked in IL, TX, EEA, UK, CH, IN: plan voiceovers
Digital Applied reports Gemini voice replication is unavailable in Illinois, Texas, the EEA, UK, Switzerland and India. A market plan using ready-made voices.
- Vrbo photo rules: 6 photos, 1024x683, no overlays
Vrbo requires at least 6 published photos at 1024 x 683 or larger, with no text or watermark overlays. Plan a compliant gallery and use AI only for promos.
Written by Sume