J cut and L cut: what they are and how to make one

A J cut plays the next shot's sound before its picture; an L cut lets a shot's sound run on under the next picture. Here's how to build each one.

5 min readSume
All posts

A J cut and an L cut are split edits: the sound and the picture change at different moments. In a J cut, the next shot's sound starts before its picture, so you hear the new scene first. In an L cut, the picture moves to the next shot while the last shot's sound keeps playing. The letters describe the shape the audio and video tracks make on an editing timeline.

The Sume facts below come from the Timeline 1.0 and Audio detach docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.

What is the difference between a J cut and an L cut?

Both start from a straight cut, where sound and picture switch together. Then the sound cut slides a second or two away from the picture cut. The direction is the only difference:

  • J cut: the sound cut comes first. The next shot's sound starts under the current picture, and the picture catches up. In an interview, the answer begins while the interviewer is still on screen.
  • L cut: the picture cut comes first. The current shot's sound carries on under the next picture, so a speaker can finish a sentence while the viewer sees the listener or the thing being described.

How do I make a J cut or an L cut with Sume?

Sume's Timeline 1.0 render keeps sound and picture apart, which is what a split edit needs. The picture is a list of video[] slots, and a picture cut happens at a slot's start. The sound is one audio spine, which you can build from up to 20 audio.parts[]: slices of audio files, each with a url, a source_in (the in-point into that file) and a duration, joined end to end. A sound cut happens where one part ends and the next begins.

  • Both clips must already be files in your workspace on media.sume.com, such as outputs of earlier Sume jobs. There is no public upload route for a file on your computer (which URLs each endpoint accepts).
  • Detach each clip's sound with POST /v1/audio-detach. Its default output is a sample-exact WAV, the format a spine wants. A render takes sound only from its spine and an optional soundtrack; in current code, each clip's own audio is dropped.
  • Give each clip one slot and one part. Keep shot B in sync: its part's source_in is its slot's source_in minus the lead for a J cut, or plus the lag for an L cut.
  • An L cut needs shot A's file to run on past its last frame on screen, and a J cut needs shot B's file to hold sound before its slot's source_in.
Shots A and B, 6 seconds each on screen, with shot B's slot at start 6 and source_in 2, and a 1.5-second split. Fields from Timeline 1.0, read 2026-09-28.
FieldStraight cutJ cutL cut
Shot A's part duration64.57.5
Shot B's part source_in20.53.5
Shot B's part duration67.54.5
Sound cut at6 s4.5 s7.5 s
Picture cut at6 s6 s6 s
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: interview-j-cut-001" \
  -d '{
    "audio": {
      "duration_seconds": 12,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_demo/a.wav", "source_in": 0, "duration": 4.5 },
        { "url": "https://media.sume.com/artifacts/artf_demo/b.wav", "source_in": 0.5, "duration": 7.5 }
      ]
    },
    "output": { "width": 1920, "height": 1080 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4", "start": 0, "duration": 6 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4", "start": 6, "duration": 6, "source_in": 2 }
    ]
  }'

What can go wrong with the sound?

The audio parts carry the split, so they are where a render gets refused or sounds wrong. Cutting away to B-roll over one continuous voice is a simpler case, covered in Add B-roll to a talking-head video.

  • The parts must add up to at least audio.duration_seconds, or the render is refused with audio_parts_shorter_than_duration.
  • Parts join in the sample domain with no silence inserted at the seams, so put each sound cut in a pause rather than in the middle of a word.
  • In current code, a render refuses parts with different channel counts (audio_parts_channel_mismatch). If one clip is mono, detach both with channels: "mono".
  • A clip with no audio track fails detach with detach_source_has_no_audio.
  • One render takes at most 20 parts. For more, join the audio first, as remove part of a video by API shows.
  • The default output is 1080×1920 (vertical). Set output.width and output.height for a horizontal edit, as the example does.

How much does a split edit cost?

Each detach is $0.01 per job and the render is $0.10 per output minute, plus a 5.5% agent fee by default. The render reserves ceil(audio.duration_seconds / 60) minutes, so the 12-second example reserves one. Before you pay, the unbilled POST /v1/timeline-1.0/plan checks the document; see validate a timeline before rendering.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume