Swapped the voice? Re-time the visuals from word timestamps

A new voice speaks at a new pace, so cuts set for the old one drift. Take each scene's start from the new take's word timestamps and re-plan the timeline.

6 min readSume
All posts

If you replace a voice-over with a different voice, do the picture cuts still line up? Not unless you re-time them. Two voices read the same script at different paces, so scene starts must come from the new take's word timestamps, not from the old timings.

This is now a routine need: Eleven v4 and MAI-Voice-2.1 both launched within days of each other (read 2026-10-04), and teams will audition them against a voice they already shipped.

Find each scene's first word

Render the new take with timestamps: { words: true }. For each scene you know the first word of its line, so search the words[] list in order and take that item's start. Searching in order, rather than from the top each time, prevents a repeated word from matching an earlier scene.

Keep a small pad before the word, such as 0.1 seconds, so the cut lands just ahead of the speech.

Build the render body

Timeline render takes audio.url, an audio.duration_seconds you measured, and video[] slots. Rules to respect: video[0].start must be 0, starts must increase, and each slot has a duration, a source_in and a fit. Each slot's duration is the next scene's start minus its own.

This script computes the slots from a list of scene openers and the new take's words.

def plan(scene_openers, words, total, pad=0.1):
    starts, i = [], 0
    for opener in scene_openers:
        while i < len(words) and words[i]["word"].lower() != opener.lower():
            i += 1
        if i == len(words):
            raise ValueError("opener not found: " + opener)
        starts.append(max(0.0, words[i]["start"] - pad))
        i += 1
    starts[0] = 0.0
    ends = starts[1:] + [total]
    return [{"start": round(s, 2), "duration": round(e - s, 2)}
            for s, e in zip(starts, ends)]

words = [{"word": "Meet", "start": 0.2}, {"word": "Then", "start": 3.9},
         {"word": "Finally", "start": 7.4}]
print(plan(["Meet", "Then", "Finally"], words, 10.0))

Preview before you pay

POST /v1/timeline-1.0/plan is an unbilled preflight that checks the same body, so run it first and fix any ordering or duration error. The render itself is $0.10 per started output minute per the guide; confirm live in GET /v1/catalog.

Add the slot fields source_in, fit and transition from your own scene plan, then render. Captions follow the same timings: send the new words to the caption job.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume