J cut and L cut: what they are and how to make one
A J cut plays the next shot's sound before its picture; an L cut lets a shot's sound run on under the next picture. Here's how to build each one.

A J cut and an L cut are split edits: the sound and the picture change at different moments. In a J cut, the next shot's sound starts before its picture, so you hear the new scene first. In an L cut, the picture moves to the next shot while the last shot's sound keeps playing. The letters describe the shape the audio and video tracks make on an editing timeline.
The Sume facts below come from the Timeline 1.0 and Audio detach docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.
What is the difference between a J cut and an L cut?
Both start from a straight cut, where sound and picture switch together. Then the sound cut slides a second or two away from the picture cut. The direction is the only difference:
- J cut: the sound cut comes first. The next shot's sound starts under the current picture, and the picture catches up. In an interview, the answer begins while the interviewer is still on screen.
- L cut: the picture cut comes first. The current shot's sound carries on under the next picture, so a speaker can finish a sentence while the viewer sees the listener or the thing being described.
How do I make a J cut or an L cut with Sume?
Sume's Timeline 1.0 render keeps sound and picture apart, which is what a split edit needs. The picture is a list of video[] slots, and a picture cut happens at a slot's start. The sound is one audio spine, which you can build from up to 20 audio.parts[]: slices of audio files, each with a url, a source_in (the in-point into that file) and a duration, joined end to end. A sound cut happens where one part ends and the next begins.
- Both clips must already be files in your workspace on
media.sume.com, such as outputs of earlier Sume jobs. There is no public upload route for a file on your computer (which URLs each endpoint accepts). - Detach each clip's sound with
POST /v1/audio-detach. Its default output is a sample-exact WAV, the format a spine wants. A render takes sound only from its spine and an optional soundtrack; in current code, each clip's own audio is dropped. - Give each clip one slot and one part. Keep shot B in sync: its part's
source_inis its slot'ssource_inminus the lead for a J cut, or plus the lag for an L cut. - An L cut needs shot A's file to run on past its last frame on screen, and a J cut needs shot B's file to hold sound before its slot's
source_in.
| Field | Straight cut | J cut | L cut |
|---|---|---|---|
Shot A's part duration | 6 | 4.5 | 7.5 |
Shot B's part source_in | 2 | 0.5 | 3.5 |
Shot B's part duration | 6 | 7.5 | 4.5 |
| Sound cut at | 6 s | 4.5 s | 7.5 s |
| Picture cut at | 6 s | 6 s | 6 s |
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: interview-j-cut-001" \
-d '{
"audio": {
"duration_seconds": 12,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/a.wav", "source_in": 0, "duration": 4.5 },
{ "url": "https://media.sume.com/artifacts/artf_demo/b.wav", "source_in": 0.5, "duration": 7.5 }
]
},
"output": { "width": 1920, "height": 1080 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/a.mp4", "start": 0, "duration": 6 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/b.mp4", "start": 6, "duration": 6, "source_in": 2 }
]
}'What can go wrong with the sound?
The audio parts carry the split, so they are where a render gets refused or sounds wrong. Cutting away to B-roll over one continuous voice is a simpler case, covered in Add B-roll to a talking-head video.
- The parts must add up to at least
audio.duration_seconds, or the render is refused withaudio_parts_shorter_than_duration. - Parts join in the sample domain with no silence inserted at the seams, so put each sound cut in a pause rather than in the middle of a word.
- In current code, a render refuses parts with different channel counts (
audio_parts_channel_mismatch). If one clip is mono, detach both withchannels: "mono". - A clip with no audio track fails detach with
detach_source_has_no_audio. - One render takes at most 20 parts. For more, join the audio first, as remove part of a video by API shows.
- The default output is 1080×1920 (vertical). Set
output.widthandoutput.heightfor a horizontal edit, as the example does.
How much does a split edit cost?
Each detach is $0.01 per job and the render is $0.10 per output minute, plus a 5.5% agent fee by default. The render reserves ceil(audio.duration_seconds / 60) minutes, so the 12-second example reserves one. Before you pay, the unbilled POST /v1/timeline-1.0/plan checks the document; see validate a timeline before rendering.
Sources
Related posts
More in Media tools
- Keyframe interval explained: I-frames, GOPs, and cuts
The keyframe interval is how far apart a video's whole-picture frames sit. It decides where copy cuts and fast seeks land. How to read and set it.
- LinkedIn ad image size: 1200×628, 1:1, and 4:5 specs
LinkedIn recommends 1200×628 for single image ads, 1200×1200 for square and 720×900 for 4:5, as JPG, PNG or GIF up to 5 MB. Carousel cards: 1080×1080.
- LinkedIn video size: dimensions, length, and file limits
A video in a LinkedIn post can be 256×144 to 4096×2304 pixels, 1:2.4 to 2.4:1, 75 KB to 5 GB and up to 15 minutes long. How to render one to fit.
- How to mix voice with background music into one file
Mix voice with background music: set the music below the voice, duck it under speech, fade it out, export one file. The Sume way: render, then detach.
Written by Sume