How to insert a video into another video
To insert a video into another video, split the main video at the insert point and play it, the new clip, then the rest, all in one render.

To insert a video into another video, split the main video at the insert point and play three pieces in order: the main video up to that point, the new clip, then the rest of the main video. Everything after the insert point moves later by the new clip's length, so the result is as long as both videos together.
The Sume facts below come from the Timeline 1.0, Audio detach and Timeline compose docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.
How do I insert a clip with Sume?
One Timeline 1.0 render does it, with no trimming first. Each video[] slot names its own source_url and source_in (the in-point into that file), so the first and third slots can both read the main video. The render's sound comes only from its audio spine and an optional soundtrack; in current code, each slot's own audio is dropped. So you rebuild the sound from the same three ranges as audio.parts[].
- Both videos must already be files in your workspace on
media.sume.com, such as outputs of earlier Sume jobs. There is no public upload route for a file on your computer (which URLs each endpoint accepts). - Detach each video's sound with
POST /v1/audio-detach. The default output is a sample-exact WAV, the format the spine wants. - Lay out three slots and three matching parts, as in the table.
audio.duration_secondsis the total length.
| Piece | `video[]` slot | `audio.parts[]` slice |
|---|---|---|
| Main video, before | start 0, duration 20, source_in 0 | Main WAV, source_in 0, duration 20 |
| New clip | start 20, duration 10, source_in 0 | Clip WAV, source_in 0, duration 10 |
| Main video, after | start 30, duration 40, source_in 20 | Main WAV, source_in 20, duration 40 |
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: insert-clip-001" \
-d '{
"audio": {
"duration_seconds": 70,
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/main.wav", "source_in": 0, "duration": 20 },
{ "url": "https://media.sume.com/artifacts/artf_demo/clip.wav", "source_in": 0, "duration": 10 },
{ "url": "https://media.sume.com/artifacts/artf_demo/main.wav", "source_in": 20, "duration": 40 }
]
},
"output": { "width": 1920, "height": 1080 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/main.mp4", "start": 0, "duration": 20 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 20, "duration": 10 },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/main.mp4", "start": 30, "duration": 40, "source_in": 20 }
]
}'What if the new clip is a different size or frame rate?
The default output is 1080×1920, so set output.width and output.height to the main video's size, as the example does. After that, a clip of another shape or frame rate is handled as in how to merge two videos of different resolutions: the slot's fit (cover by default) decides how it fills the frame, and a clip at another rate repeats or drops frames, with the warning output_fps_resamples_sources.
Can I show the second video on top of the first instead?
Not with Sume's media tools. An insert plays the clips one after another, and timeline compose, the tool for two sources on screen at once, takes one still and one video, not two videos. If the new clip should play over the main video's sound without pausing it, that is a cutaway: see Add B-roll to a talking-head video. To insert a still image, see how to add an image to a video at a specific time.
What does it cost, and what are the limits?
Each detach is $0.01 per job and the render is $0.10 per output minute, plus a 5.5% agent fee by default. The render reserves ceil(audio.duration_seconds / 60) minutes, so the 70-second example reserves two. Check the document first with the unbilled POST /v1/timeline-1.0/plan, covered in validate a timeline before rendering.
- The output runs 1–1,800 seconds, and each slot lasts at least 0.2 seconds.
- In current code, detach and the render refuse any source file over 300 MiB with
source_too_large. - One render takes up to 20 audio parts, and the parts must add up to at least
audio.duration_seconds. - One detach writes at most 900 seconds. A longer main video needs two detaches with
range, and each part'ssource_inthen counts from the start of its own file. - A clip with no audio track fails detach with
detach_source_has_no_audio. A free inspect withframes: falseshowsprobe.has_audiofirst. - In current code, parts with different channel counts are refused (
audio_parts_channel_mismatch). If one video is mono, detach both withchannels: "mono".
Sources
Related posts
More in Media tools
- J cut and L cut: what they are and how to make one
A J cut plays the next shot's sound before its picture; an L cut lets a shot's sound run on under the next picture. Here's how to build each one.
- Keyframe interval explained: I-frames, GOPs, and cuts
The keyframe interval is how far apart a video's whole-picture frames sit. It decides where copy cuts and fast seeks land. How to read and set it.
- LinkedIn ad image size: 1200×628, 1:1, and 4:5 specs
LinkedIn recommends 1200×628 for single image ads, 1200×1200 for square and 720×900 for 4:5, as JPG, PNG or GIF up to 5 MB. Carousel cards: 1080×1080.
- LinkedIn video size: dimensions, length, and file limits
A video in a LinkedIn post can be 256×144 to 4096×2304 pixels, 1:2.4 to 2.4:1, 75 KB to 5 GB and up to 15 minutes long. How to render one to fit.
Written by Sume