Does the ad voice start in the first second? STT first-word check

Run recorded ads through Sume STT and read words[0].start. Flag any clip where speech starts after 1.0 s. A 20-ad batch is about 10 cents of usage.

4 min readSume
All posts

Short ads live and die on the first seconds. A clip that opens on 1.8 seconds of room tone before the first word wastes the hook. A person can see that in an editor. A batch of 60 exports from five editors needs a script. Sume STT 1.0 gives you the time of the first spoken word as words[0].start, which is all the check needs.

Rules of the check

Read words and skip entries that are not type == "word". The first real word's start is the delay. Pick a limit that fits your format, such as 1.0 seconds, and flag anything over it. For clips that start with music and no voice on purpose, set a higher limit per campaign. Report the delay in the file name or a sheet, so the editor can trim without opening the file.

Code

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

def first_word(url):
    res = run("/stt-1.0/transcribe", {"audio_url": url, "duration_seconds": 30})
    for w in res["words"]:
        if w.get("type") == "word":
            return w["start"]
    return None

LIMIT = 1.0
for url in open("ads.txt").read().split():
    t = first_word(url)
    flag = "NO SPEECH" if t is None else ("LATE" if t > LIMIT else "ok")
    print(flag, t, url)

What it can and cannot tell you

  • It measures when speech starts, not whether the hook is good.
  • It works on the audio the URL points to. For a finished video, extract the audio first (Sume has POST /v1/audio-detach, wav by default) and send that URL; the audio must be a public HTTPS URL for STT.
  • A cough or breath before the first word can be transcribed as a word or not. Spot-check five flagged clips.
  • If the clip has no speech, words can be empty, which the code reports as NO SPEECH.

Cost

STT 1.0 is about $0.01 per audio minute. A batch of 20 ads of 30 seconds is 10 minutes of audio, about 10 cents of usage. Set duration_seconds to a number just over the clip length so the reservation matches the work. When the check finds a late start, the fix is a trim, not a new recording. See the lead-in trim post for the cut.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume