Two-camera podcast clip for Shorts: alternate angles with source_in

Cut between two camera files against one audio spine in Timeline 1.0: equal on-spine starts and source_in keep both angles in sync. Python builder included.

6 min readSume
All posts

The pattern: one spine, two angles, same clock

A two-camera interview becomes a Short by cutting between the two angles while the conversation audio plays once, underneath. Timeline 1.0 is built the same way: one audio spine and an ordered video[] of slots, where each slot says where it starts on the spine (start), how long it stays (duration) and where in its own file it starts reading (source_in). If both cameras began recording together and you cut at on-spine second 12, a slot with start: 12 and source_in: 12 shows what camera A or B saw at that moment, in sync with the voice.

That makes the whole edit a list of numbers, and a list of numbers is something a script can generate. The October platform roundup says YouTube's Shorts series, with seasons and episodes, began rolling out on 23 September. A weekly interview series cut into Shorts is exactly the kind of repeat job where a generator beats a manual edit.

Constraints that shape the cut

These come from the Timeline 1.0 program table; the generator below respects all of them.

Timeline slot rules that matter for an angle-switching cut (read 2026-10-06)
RuleValueWhy it matters here
video[0].startMust be 0The first angle starts with the spine
Later startsMust increaseEach cut moves forward on the spine
Slot durationAt least 0.2 sNo sliver slots at the end
CoverageEnds at most 0.5 s before the spine endThe last slot must reach the end
Slots per render1 to 200About 33 slots for a 200 s clip at 6 s cuts
SourcesMust be media.sume.com artifacts of your workspaceImport both cameras first

A body builder

multicam alternates the two files every cut_every seconds, sets source_in equal to the on-spine start so the angles stay synced, and folds a would-be sliver into the last slot. The demo builds a 47-second clip and prints the slot count and the end time, with no network. Post the body to /v1/timeline-1.0/plan first and then to /v1/timeline-1.0/render.

import json

def multicam(cam_a, cam_b, voice, total, cut_every=6):
    video, t, i = [], 0, 0
    while t < total:
        d = min(cut_every, total - t)
        if total - (t + d) < 0.2:   # never leave a sliver slot behind
            d = total - t
        video.append({"source_url": (cam_a, cam_b)[i % 2], "start": t,
                      "duration": d, "source_in": t})
        t, i = t + d, i + 1
    return {"audio": {"url": voice, "duration_seconds": total}, "video": video,
            "output": {"width": 1080, "height": 1920, "fps": 30}}

body = multicam("https://media.sume.com/artifacts/artf_demo/cam-a.mp4",
                "https://media.sume.com/artifacts/artf_demo/cam-b.mp4",
                "https://media.sume.com/artifacts/artf_demo/mix.wav", 47)
print(len(body["video"]), "slots, last ends at",
      body["video"][-1]["start"] + body["video"][-1]["duration"])
print(json.dumps(body["video"][:2], indent=1))

Audio and sync: the parts the script cannot do

The spine is a Sume-hosted audio file. If your good audio is the track of one camera, audio detach turns it into a durable audio file you can use as the spine; its output is capped at 900 seconds. If the cameras did not start together, the equal source_in trick breaks, and you need a per-camera offset: add it to source_in for that camera only. Sume does not detect sync for you. A clap or a slate at the start makes measuring the offset easy.

A cut every six seconds is a default for a script, not an editing recommendation. Real conversations cut on speaker changes, and finding those needs a transcript with timings. Video inspect can transcribe the audio at $0.01 per audio minute, and if the transcript carries timings you can use them as cut points instead of a fixed interval; check the response shape in the docs first.

Framing for vertical

The default output is 1080 by 1920. A landscape camera file will be fitted with fit, which defaults to cover and crops the sides; contain letterboxes, and blur fills the bars with a blurred copy. For two people side by side in a wide shot, cover can crop one of them out. Decide fit per camera, and look at the first plan-approved render before you run the whole season.

Check the result before you scale up. Render one clip, then scrub the cuts against the voice: the lip movement on each cut should match the audio, with no drift. If the angles are off by a fixed amount, the cameras were not started together, so add that fixed offset to the source_in of the later camera and render again. If the drift grows over the clip, the cameras ran at slightly different clocks, and a cut-based edit cannot fix that. Re-sync at the source.

Use the unbilled plan for the first look. It reports duration_seconds, segment_count and billable_minutes, so you will see how many slots a 47-second clip produced and what the render will bill before you pay for it. Renders are billed at $0.10 per whole output minute under the documented rate.

If you would rather have fewer, longer cuts, raise cut_every and the slot count drops with it. Fewer slots also means fewer places for a mistake, and a single-take clip of one camera with an occasional cutaway is often a better Short than a rapid ping-pong between two heads.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume