Two-camera podcast clip for Shorts: alternate angles with source_in
Cut between two camera files against one audio spine in Timeline 1.0: equal on-spine starts and source_in keep both angles in sync. Python builder included.

The pattern: one spine, two angles, same clock
A two-camera interview becomes a Short by cutting between the two angles while the conversation audio plays once, underneath. Timeline 1.0 is built the same way: one audio spine and an ordered video[] of slots, where each slot says where it starts on the spine (start), how long it stays (duration) and where in its own file it starts reading (source_in). If both cameras began recording together and you cut at on-spine second 12, a slot with start: 12 and source_in: 12 shows what camera A or B saw at that moment, in sync with the voice.
That makes the whole edit a list of numbers, and a list of numbers is something a script can generate. The October platform roundup says YouTube's Shorts series, with seasons and episodes, began rolling out on 23 September. A weekly interview series cut into Shorts is exactly the kind of repeat job where a generator beats a manual edit.
Constraints that shape the cut
These come from the Timeline 1.0 program table; the generator below respects all of them.
| Rule | Value | Why it matters here |
|---|---|---|
| video[0].start | Must be 0 | The first angle starts with the spine |
| Later starts | Must increase | Each cut moves forward on the spine |
| Slot duration | At least 0.2 s | No sliver slots at the end |
| Coverage | Ends at most 0.5 s before the spine end | The last slot must reach the end |
| Slots per render | 1 to 200 | About 33 slots for a 200 s clip at 6 s cuts |
| Sources | Must be media.sume.com artifacts of your workspace | Import both cameras first |
A body builder
multicam alternates the two files every cut_every seconds, sets source_in equal to the on-spine start so the angles stay synced, and folds a would-be sliver into the last slot. The demo builds a 47-second clip and prints the slot count and the end time, with no network. Post the body to /v1/timeline-1.0/plan first and then to /v1/timeline-1.0/render.
import json
def multicam(cam_a, cam_b, voice, total, cut_every=6):
video, t, i = [], 0, 0
while t < total:
d = min(cut_every, total - t)
if total - (t + d) < 0.2: # never leave a sliver slot behind
d = total - t
video.append({"source_url": (cam_a, cam_b)[i % 2], "start": t,
"duration": d, "source_in": t})
t, i = t + d, i + 1
return {"audio": {"url": voice, "duration_seconds": total}, "video": video,
"output": {"width": 1080, "height": 1920, "fps": 30}}
body = multicam("https://media.sume.com/artifacts/artf_demo/cam-a.mp4",
"https://media.sume.com/artifacts/artf_demo/cam-b.mp4",
"https://media.sume.com/artifacts/artf_demo/mix.wav", 47)
print(len(body["video"]), "slots, last ends at",
body["video"][-1]["start"] + body["video"][-1]["duration"])
print(json.dumps(body["video"][:2], indent=1))
Audio and sync: the parts the script cannot do
The spine is a Sume-hosted audio file. If your good audio is the track of one camera, audio detach turns it into a durable audio file you can use as the spine; its output is capped at 900 seconds. If the cameras did not start together, the equal source_in trick breaks, and you need a per-camera offset: add it to source_in for that camera only. Sume does not detect sync for you. A clap or a slate at the start makes measuring the offset easy.
A cut every six seconds is a default for a script, not an editing recommendation. Real conversations cut on speaker changes, and finding those needs a transcript with timings. Video inspect can transcribe the audio at $0.01 per audio minute, and if the transcript carries timings you can use them as cut points instead of a fixed interval; check the response shape in the docs first.
Framing for vertical
The default output is 1080 by 1920. A landscape camera file will be fitted with fit, which defaults to cover and crops the sides; contain letterboxes, and blur fills the bars with a blurred copy. For two people side by side in a wide shot, cover can crop one of them out. Decide fit per camera, and look at the first plan-approved render before you run the whole season.
Check the result before you scale up. Render one clip, then scrub the cuts against the voice: the lip movement on each cut should match the audio, with no drift. If the angles are off by a fixed amount, the cameras were not started together, so add that fixed offset to the source_in of the later camera and render again. If the drift grows over the clip, the cameras ran at slightly different clocks, and a cut-based edit cannot fix that. Re-sync at the source.
Use the unbilled plan for the first look. It reports duration_seconds, segment_count and billable_minutes, so you will see how many slots a 47-second clip produced and what the render will bill before you pay for it. Renders are billed at $0.10 per whole output minute under the documented rate.
If you would rather have fewer, longer cuts, raise cut_every and the slot count drops with it. Fewer slots also means fewer places for a mistake, and a single-take clip of one camera with an occasional cutaway is often a better Short than a rapid ping-pong between two heads.
Sources
Related posts
More in Use cases
- Two voices, one conversation: Sume TTS jobs joined by concat
Make a two-person dialogue file with Sume: one TTS job per line using two avatar voices, then one Timeline audio concat. Cost, code and the gap.
- UGC-style ad: test five hooks on one body with Timeline plans
Join five 3-second hooks to one 12-second body clip with Timeline 1.0. Plan each cut unbilled, then render the winners at $0.10 a minute.
- UGC-style ad: a music bed that ducks under the voice
Add a looped music bed to a UGC-style ad and lower it under the voice with soundtrack.duck_db in a Timeline 1.0 render. Needs a real voice spine.
- Video ad end card: add a fade-in and fade-out with Timeline
Close an ad with a 1-second fade-out and open it with a 0.5-second fade-in using output.fade_in_seconds and fade_out_seconds in a Timeline 1.0 render.
Written by Sume