How to insert a video into another video

To insert a video into another video, split the main video at the insert point and play it, the new clip, then the rest, all in one render.

5 min readSume
All posts

To insert a video into another video, split the main video at the insert point and play three pieces in order: the main video up to that point, the new clip, then the rest of the main video. Everything after the insert point moves later by the new clip's length, so the result is as long as both videos together.

The Sume facts below come from the Timeline 1.0, Audio detach and Timeline compose docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.

How do I insert a clip with Sume?

One Timeline 1.0 render does it, with no trimming first. Each video[] slot names its own source_url and source_in (the in-point into that file), so the first and third slots can both read the main video. The render's sound comes only from its audio spine and an optional soundtrack; in current code, each slot's own audio is dropped. So you rebuild the sound from the same three ranges as audio.parts[].

  • Both videos must already be files in your workspace on media.sume.com, such as outputs of earlier Sume jobs. There is no public upload route for a file on your computer (which URLs each endpoint accepts).
  • Detach each video's sound with POST /v1/audio-detach. The default output is a sample-exact WAV, the format the spine wants.
  • Lay out three slots and three matching parts, as in the table. audio.duration_seconds is the total length.
A 10-second clip inserted at 0:20 of a 60-second video; the output runs 70 seconds. Fields from Timeline 1.0, read 2026-09-28.
Piece`video[]` slot`audio.parts[]` slice
Main video, beforestart 0, duration 20, source_in 0Main WAV, source_in 0, duration 20
New clipstart 20, duration 10, source_in 0Clip WAV, source_in 0, duration 10
Main video, afterstart 30, duration 40, source_in 20Main WAV, source_in 20, duration 40
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: insert-clip-001" \
  -d '{
    "audio": {
      "duration_seconds": 70,
      "parts": [
        { "url": "https://media.sume.com/artifacts/artf_demo/main.wav", "source_in": 0, "duration": 20 },
        { "url": "https://media.sume.com/artifacts/artf_demo/clip.wav", "source_in": 0, "duration": 10 },
        { "url": "https://media.sume.com/artifacts/artf_demo/main.wav", "source_in": 20, "duration": 40 }
      ]
    },
    "output": { "width": 1920, "height": 1080 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/main.mp4", "start": 0, "duration": 20 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 20, "duration": 10 },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/main.mp4", "start": 30, "duration": 40, "source_in": 20 }
    ]
  }'

What if the new clip is a different size or frame rate?

The default output is 1080×1920, so set output.width and output.height to the main video's size, as the example does. After that, a clip of another shape or frame rate is handled as in how to merge two videos of different resolutions: the slot's fit (cover by default) decides how it fills the frame, and a clip at another rate repeats or drops frames, with the warning output_fps_resamples_sources.

Can I show the second video on top of the first instead?

Not with Sume's media tools. An insert plays the clips one after another, and timeline compose, the tool for two sources on screen at once, takes one still and one video, not two videos. If the new clip should play over the main video's sound without pausing it, that is a cutaway: see Add B-roll to a talking-head video. To insert a still image, see how to add an image to a video at a specific time.

What does it cost, and what are the limits?

Each detach is $0.01 per job and the render is $0.10 per output minute, plus a 5.5% agent fee by default. The render reserves ceil(audio.duration_seconds / 60) minutes, so the 70-second example reserves two. Check the document first with the unbilled POST /v1/timeline-1.0/plan, covered in validate a timeline before rendering.

  • The output runs 1–1,800 seconds, and each slot lasts at least 0.2 seconds.
  • In current code, detach and the render refuse any source file over 300 MiB with source_too_large.
  • One render takes up to 20 audio parts, and the parts must add up to at least audio.duration_seconds.
  • One detach writes at most 900 seconds. A longer main video needs two detaches with range, and each part's source_in then counts from the start of its own file.
  • A clip with no audio track fails detach with detach_source_has_no_audio. A free inspect with frames: false shows probe.has_audio first.
  • In current code, parts with different channel counts are refused (audio_parts_channel_mismatch). If one video is mono, detach both with channels: "mono".

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume