Does the ad voice start in the first second? STT first-word check
Run recorded ads through Sume STT and read words[0].start. Flag any clip where speech starts after 1.0 s. A 20-ad batch is about 10 cents of usage.

Short ads live and die on the first seconds. A clip that opens on 1.8 seconds of room tone before the first word wastes the hook. A person can see that in an editor. A batch of 60 exports from five editors needs a script. Sume STT 1.0 gives you the time of the first spoken word as words[0].start, which is all the check needs.
Rules of the check
Read words and skip entries that are not type == "word". The first real word's start is the delay. Pick a limit that fits your format, such as 1.0 seconds, and flag anything over it. For clips that start with music and no voice on purpose, set a higher limit per campaign. Report the delay in the file name or a sheet, so the editor can trim without opening the file.
Code
import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = {**H, **({"Idempotency-Key": key} if key else {})}
r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
r.raise_for_status()
job = r.json()["data"]["job"]["id"]
while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
time.sleep(3)
res = requests.get(f"{B}/jobs/{job}/result", headers=H)
res.raise_for_status()
return res.json()["data"]["result"]
def first_word(url):
res = run("/stt-1.0/transcribe", {"audio_url": url, "duration_seconds": 30})
for w in res["words"]:
if w.get("type") == "word":
return w["start"]
return None
LIMIT = 1.0
for url in open("ads.txt").read().split():
t = first_word(url)
flag = "NO SPEECH" if t is None else ("LATE" if t > LIMIT else "ok")
print(flag, t, url)What it can and cannot tell you
- It measures when speech starts, not whether the hook is good.
- It works on the audio the URL points to. For a finished video, extract the audio first (Sume has
POST /v1/audio-detach, wav by default) and send that URL; the audio must be a public HTTPS URL for STT. - A cough or breath before the first word can be transcribed as a word or not. Spot-check five flagged clips.
- If the clip has no speech,
wordscan be empty, which the code reports as NO SPEECH.
Cost
STT 1.0 is about $0.01 per audio minute. A batch of 20 ads of 30 seconds is 10 minutes of audio, about 10 cents of usage. Set duration_seconds to a number just over the clip length so the reservation matches the work. When the check finds a late start, the fix is a trim, not a new recording. See the lead-in trim post for the cut.
Sources
Related posts
More in Use cases
- IT helpdesk password reset video: 30-second AI avatar clip and cost
A reusable helpdesk how-to clip from one Sume avatar: a 30-second password reset script, inline captions, and the cost per tier. One render, many viewers.
- Jazz cafe music for a 30-minute screen loop: about $3 on Sume
One $0.125 jazz track plus a 30-minute Timeline render with soundtrack loop is $3.125. How to brief it, loop it, and fade it, with the 1,800 s cap.
- Jewelry macro in one take: Seedance 2.5 prompt, length and price
A macro-lens prompt for a ring, watch or pendant on Seedance 2.5: slow rotation, single light source, 6 to 12 seconds, and the Sume price at 720p and 1080p.
- K-pop style instrumental for a product video: a brief with no artist
Brief Lyria 3.5 on Sume for a K-pop style instrumental by tempo, synths, and drops, not artist names. $0.125 per take; Timeline mixes it at $0.10 a minute.
Written by Sume