Swapped the voice? Re-time the visuals from word timestamps
A new voice speaks at a new pace, so cuts set for the old one drift. Take each scene's start from the new take's word timestamps and re-plan the timeline.

If you replace a voice-over with a different voice, do the picture cuts still line up? Not unless you re-time them. Two voices read the same script at different paces, so scene starts must come from the new take's word timestamps, not from the old timings.
This is now a routine need: Eleven v4 and MAI-Voice-2.1 both launched within days of each other (read 2026-10-04), and teams will audition them against a voice they already shipped.
Find each scene's first word
Render the new take with timestamps: { words: true }. For each scene you know the first word of its line, so search the words[] list in order and take that item's start. Searching in order, rather than from the top each time, prevents a repeated word from matching an earlier scene.
Keep a small pad before the word, such as 0.1 seconds, so the cut lands just ahead of the speech.
Build the render body
Timeline render takes audio.url, an audio.duration_seconds you measured, and video[] slots. Rules to respect: video[0].start must be 0, starts must increase, and each slot has a duration, a source_in and a fit. Each slot's duration is the next scene's start minus its own.
This script computes the slots from a list of scene openers and the new take's words.
def plan(scene_openers, words, total, pad=0.1):
starts, i = [], 0
for opener in scene_openers:
while i < len(words) and words[i]["word"].lower() != opener.lower():
i += 1
if i == len(words):
raise ValueError("opener not found: " + opener)
starts.append(max(0.0, words[i]["start"] - pad))
i += 1
starts[0] = 0.0
ends = starts[1:] + [total]
return [{"start": round(s, 2), "duration": round(e - s, 2)}
for s, e in zip(starts, ends)]
words = [{"word": "Meet", "start": 0.2}, {"word": "Then", "start": 3.9},
{"word": "Finally", "start": 7.4}]
print(plan(["Meet", "Then", "Finally"], words, 10.0))
Preview before you pay
POST /v1/timeline-1.0/plan is an unbilled preflight that checks the same body, so run it first and fix any ordering or duration error. The render itself is $0.10 per started output minute per the guide; confirm live in GET /v1/catalog.
Add the slot fields source_in, fit and transition from your own scene plan, then render. Captions follow the same timings: send the new words to the caption job.
Sources
Related posts
More in Media tools
- Reference ingest coverage: how many frames OCR read
The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.
- Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.
- Reference ingest shot record: cut types, palette, luma, motion
Each reference-ingest shot carries cut_out type (hard, gradual, end), palette, luma, contrast and a motion class. Turn them into shot length and pacing numbers.
- Reference ingest source block: rotation, vfr and aspect
Before you remake a reference, read the manifest source block: display_aspect, rotation, fps, vfr, codec and has_audio_track. What each field should change.
Written by Sume