Show the product name the moment the voice says it: STT word times

Find when a voiceover says a word with Sume STT words[], then use that second as the next timeline video[].start. Offline Python script included.

5 min readSume
All posts

Short answer

To put a visual on screen at the instant a spoken word lands, transcribe the voiceover, find the word in words[], and use its start as the start of the next video[] slot on a Sume timeline. The word times come back with every transcript, and the timeline accepts declared starts as authoritative.

The pieces are documented: the STT result has words[] of {word, start, end}, and the Timeline 1.0 docs say video[0].start must be 0, later starts must increase, and declared starts are authoritative.

How the timeline reads starts

Timeline slots are time-based, so the first slot always begins at 0 and runs until the second slot's start. If the product name is spoken at 7.42 seconds, the shot before it runs from 0 to 7.42 and the product shot begins at 7.42. A slot's duration must be at least 0.2 seconds, and coverage can stop at most 0.5 seconds before the end of the audio.

The script below takes STT-shaped words, finds the first match of a target word, and builds two slots. The sample numbers are invented to show the shape.

import re

words = [
    {"word": "Meet", "start": 0.40, "end": 0.62},
    {"word": "the", "start": 0.62, "end": 0.70},
    {"word": "Lumio,", "start": 0.70, "end": 1.18},
    {"word": "a", "start": 1.30, "end": 1.36},
    {"word": "lamp", "start": 1.36, "end": 1.80},
]

def norm(w):
    return re.sub(r"[^\w]", "", w).lower()

def start_of(target):
    for w in words:
        if norm(w["word"]) == norm(target):
            return w["start"]
    return None

t = start_of("Lumio")
if t is None or t < 0.2:
    raise SystemExit("word not found, or too early for a first slot")

video = [
    {"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": round(t, 2)},
    {"source_url": "https://media.sume.com/artifacts/artf_demo/product.mp4", "start": round(t, 2), "duration": 4.0},
]
print(video)

Matching the word

Normalizing punctuation matters: the transcript can return a word with a trailing comma, as in the sample, and a plain equality check would miss it. Matching on a second occurrence, or a phrase, needs a loop over the list, since a name can be spoken twice.

Lead time and checking

Add a lead of a few tenths of a second if the visual should appear just before the word, and keep the lead the same across a series. Word times are model output, not frame-exact, so check a sample render. A transition on the second slot adds its own duration, and the compiler compensates for xfade rather than shifting your declared starts.

Cost

Speech-to-text is $0.01 per audio minute at the public rate. A timeline render is $0.10 per output minute, rounded up. Check the live prices before budgeting.

Several names in one voiceover

The same approach extends to a list of names. Run start_of for each target in the order they are spoken, keep only the hits, sort by time, and turn each into one slot. The slot before a hit ends where the next hit starts, so you only compute starts; each duration is the difference between neighbors, and the last slot runs to the end of the audio.

Check two rules before you submit. Each start must be greater than the one before it, and each duration must be at least 0.2 seconds, so two names spoken closer together than that cannot each get a slot. Merge them into one slot or drop the second.

The unbilled POST /v1/timeline-1.0/plan route compiles the document without creating a job, which makes it a cheap place to catch a bad start before the render.

  • Sort hits by time and drop any pair closer than 0.2 s.
  • Last slot ends at the audio length, within the 0.5 s allowance.
  • Plan first, render second.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume