Show the product name the moment the voice says it: STT word times
Find when a voiceover says a word with Sume STT words[], then use that second as the next timeline video[].start. Offline Python script included.

Short answer
To put a visual on screen at the instant a spoken word lands, transcribe the voiceover, find the word in words[], and use its start as the start of the next video[] slot on a Sume timeline. The word times come back with every transcript, and the timeline accepts declared starts as authoritative.
The pieces are documented: the STT result has words[] of {word, start, end}, and the Timeline 1.0 docs say video[0].start must be 0, later starts must increase, and declared starts are authoritative.
How the timeline reads starts
Timeline slots are time-based, so the first slot always begins at 0 and runs until the second slot's start. If the product name is spoken at 7.42 seconds, the shot before it runs from 0 to 7.42 and the product shot begins at 7.42. A slot's duration must be at least 0.2 seconds, and coverage can stop at most 0.5 seconds before the end of the audio.
The script below takes STT-shaped words, finds the first match of a target word, and builds two slots. The sample numbers are invented to show the shape.
import re
words = [
{"word": "Meet", "start": 0.40, "end": 0.62},
{"word": "the", "start": 0.62, "end": 0.70},
{"word": "Lumio,", "start": 0.70, "end": 1.18},
{"word": "a", "start": 1.30, "end": 1.36},
{"word": "lamp", "start": 1.36, "end": 1.80},
]
def norm(w):
return re.sub(r"[^\w]", "", w).lower()
def start_of(target):
for w in words:
if norm(w["word"]) == norm(target):
return w["start"]
return None
t = start_of("Lumio")
if t is None or t < 0.2:
raise SystemExit("word not found, or too early for a first slot")
video = [
{"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": round(t, 2)},
{"source_url": "https://media.sume.com/artifacts/artf_demo/product.mp4", "start": round(t, 2), "duration": 4.0},
]
print(video)Matching the word
Normalizing punctuation matters: the transcript can return a word with a trailing comma, as in the sample, and a plain equality check would miss it. Matching on a second occurrence, or a phrase, needs a loop over the list, since a name can be spoken twice.
Lead time and checking
Add a lead of a few tenths of a second if the visual should appear just before the word, and keep the lead the same across a series. Word times are model output, not frame-exact, so check a sample render. A transition on the second slot adds its own duration, and the compiler compensates for xfade rather than shifting your declared starts.
Cost
Speech-to-text is $0.01 per audio minute at the public rate. A timeline render is $0.10 per output minute, rounded up. Check the live prices before budgeting.
Several names in one voiceover
The same approach extends to a list of names. Run start_of for each target in the order they are spoken, keep only the hits, sort by time, and turn each into one slot. The slot before a hit ends where the next hit starts, so you only compute starts; each duration is the difference between neighbors, and the last slot runs to the end of the audio.
Check two rules before you submit. Each start must be greater than the one before it, and each duration must be at least 0.2 seconds, so two names spoken closer together than that cannot each get a slot. Merge them into one slot or drop the second.
The unbilled POST /v1/timeline-1.0/plan route compiles the document without creating a job, which makes it a cheap place to catch a bad start before the render.
- Sort hits by time and drop any pair closer than 0.2 s.
- Last slot ends at the audio length, within the 0.5 s allowance.
- Plan first, render second.
Sources
Related posts
More in Developers
- Nightly video batch after Sora: 6, 24, 48 or 120 jobs by Sume plan
A cron that submitted 30 Sora renders at once needs a new ceiling. Sume accepts 6 jobs on Free, 24 on Pro, 48 on Startup, 120 on Scale before queue_full.
- Node 22 script: create a Sume bulk queue and poll it to the end
Dependency-free Node 22 ESM script: POST a bulk queue from items.json, back off the poll, survive 429 and 503, and exit non-zero when any item failed.
- Omni draft grid: four 360p variants, then one final. What it costs
Google's Draft Room idea, run through the Sume API: four 8-second 360p drafts that change one thing each, then a 1080p final. Total $2.70, with a script.
- Pandas DataFrame to a Sume bulk queue in 100-row chunks
Turn a product DataFrame into Sume Format bulk queues: one item per row, 100 rows per queue, a stable key per chunk, and SKU order saved beside each queue id.
Written by Sume