FFmpeg merge audio and video: -map, copy, and -shortest

Merge audio and video with FFmpeg: map the video from one file and the audio from the other, copy the video, and add -shortest to match lengths.

5 min readSume
All posts

To merge audio and video with FFmpeg, give both files as inputs and map the video stream from the first and the audio stream from the second: ffmpeg -i video.mp4 -i audio.wav -map 0:v -map 1:a -c:v copy -c:a aac -shortest output.mp4. The video is copied as it is, the audio is encoded to AAC, and -shortest ends the file when the shorter of the two ends.

The FFmpeg facts come from its ffmpeg, codecs, and filters documentation, and the Sume facts from the Timeline 1.0 docs and the Sume API reference, all read on 2026-09-28. Anything described as current behavior is read from Sume's code.

What does each part of the command do?

The two -map options do the work. FFmpeg's own docs combine streams from two files the same way, ffmpeg -i INPUT0.mkv -i INPUT1.aac -map 0:0 -map 1:0 -c copy OUTPUT.mp4, so when the audio is already AAC, -c copy copies both streams and nothing is re-encoded.

From the ffmpeg and codecs documentation, read 2026-09-28.
OptionWhat it does
-i video.mp4 -i audio.wavTwo inputs, numbered 0 and 1 in the order given
-map 0:vTakes the video from input 0. Any -map turns off FFmpeg's default stream selection for that output.
-map 1:aTakes the audio from input 1, so the video file's own sound is left out
-c:v copyCopies the video stream without re-encoding it: no quality loss
-c:a aacEncodes the audio with FFmpeg's native AAC encoder, at 128 kbps unless you set -b:a
-shortestFinishes encoding when the shortest output stream ends

What happens when the audio and video lengths differ?

You choose, with one option each:

  • -shortest ends the file with the shorter stream: a long song is cut where the video ends, and a short voiceover cuts the video where it ends.
  • To keep the whole video and pad a shorter track with silence, add FFmpeg's apad filter with -shortest. Its docs describe exactly this pairing, which extends the audio to the video's length: ffmpeg -i video.mp4 -i audio.wav -map 0:v -map 1:a -c:v copy -c:a aac -af apad -shortest output.mp4.
  • To end at an exact length instead, set -t, which stops writing the output after that duration.

How do I keep the video's own sound and add the new audio?

Mix the two tracks with the amix filter instead of replacing one. duration=first ends the mix with the first input, here the video's own sound:

  • weights sets each input's share, as in FFmpeg's own example that gives music a quarter of the vocals' weight. That example also sets normalize=0: by default, amix scales its inputs instead of only summing them.
  • The video file must have an audio stream for [0:a] to exist. For the Sume version of a voiceover over a clip's sound, see how to add a voiceover to a video.
ffmpeg -i video.mp4 -i voice.wav -filter_complex \
'[0:a][1:a]amix=inputs=2:duration=first[a]' \
-map 0:v -map '[a]' -c:v copy -c:a aac output.mp4

How do I merge audio and video with the Sume API?

Render the clip over the new track with Timeline 1.0: the audio file becomes the render's spine, audio.url, and audio.duration_seconds sets the output length. In current code the render re-encodes the video, unlike -c:v copy, and leaves the clip's own sound out. Set output to the clip's size, since the default frame is 1080×1920. Replace or remove the audio in a video shows the request.

Where -shortest stops at the shorter stream, the render always follows the audio: the API reference says the output lasts exactly audio.duration_seconds.

  • Video shorter than the audio: in current code, slots that end early hold their last frame to the end and warn video_coverage_shorter_than_audio, and a slot longer than its source file is filled by render.pad_mode, which loops short sources by default. Video freezes but audio continues covers both cases.
  • Video longer than the audio: end the last slot at the spine. In current code a slot that ends more than 0.5 seconds past it is refused with invalid_segment_timing.
  • Both files must already be your workspace's media.sume.com files, such as an audio detach WAV or a TTS master for the spine; which URLs each endpoint accepts explains the rule. In current code a source file over 300 MiB is refused with source_too_large.
  • A render is listed at $0.10 per output minute on API pricing and reserves ceil(audio.duration_seconds / 60) minutes, plus a 5.5% agent fee by default.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume