FFmpeg merge audio and video: -map, copy, and -shortest
Merge audio and video with FFmpeg: map the video from one file and the audio from the other, copy the video, and add -shortest to match lengths.

To merge audio and video with FFmpeg, give both files as inputs and map the video stream from the first and the audio stream from the second: ffmpeg -i video.mp4 -i audio.wav -map 0:v -map 1:a -c:v copy -c:a aac -shortest output.mp4. The video is copied as it is, the audio is encoded to AAC, and -shortest ends the file when the shorter of the two ends.
The FFmpeg facts come from its ffmpeg, codecs, and filters documentation, and the Sume facts from the Timeline 1.0 docs and the Sume API reference, all read on 2026-09-28. Anything described as current behavior is read from Sume's code.
What does each part of the command do?
The two -map options do the work. FFmpeg's own docs combine streams from two files the same way, ffmpeg -i INPUT0.mkv -i INPUT1.aac -map 0:0 -map 1:0 -c copy OUTPUT.mp4, so when the audio is already AAC, -c copy copies both streams and nothing is re-encoded.
| Option | What it does |
|---|---|
-i video.mp4 -i audio.wav | Two inputs, numbered 0 and 1 in the order given |
-map 0:v | Takes the video from input 0. Any -map turns off FFmpeg's default stream selection for that output. |
-map 1:a | Takes the audio from input 1, so the video file's own sound is left out |
-c:v copy | Copies the video stream without re-encoding it: no quality loss |
-c:a aac | Encodes the audio with FFmpeg's native AAC encoder, at 128 kbps unless you set -b:a |
-shortest | Finishes encoding when the shortest output stream ends |
What happens when the audio and video lengths differ?
You choose, with one option each:
-shortestends the file with the shorter stream: a long song is cut where the video ends, and a short voiceover cuts the video where it ends.- To keep the whole video and pad a shorter track with silence, add FFmpeg's
apadfilter with-shortest. Its docs describe exactly this pairing, which extends the audio to the video's length:ffmpeg -i video.mp4 -i audio.wav -map 0:v -map 1:a -c:v copy -c:a aac -af apad -shortest output.mp4. - To end at an exact length instead, set
-t, which stops writing the output after that duration.
How do I keep the video's own sound and add the new audio?
Mix the two tracks with the amix filter instead of replacing one. duration=first ends the mix with the first input, here the video's own sound:
weightssets each input's share, as in FFmpeg's own example that gives music a quarter of the vocals' weight. That example also setsnormalize=0: by default, amix scales its inputs instead of only summing them.- The video file must have an audio stream for
[0:a]to exist. For the Sume version of a voiceover over a clip's sound, see how to add a voiceover to a video.
ffmpeg -i video.mp4 -i voice.wav -filter_complex \
'[0:a][1:a]amix=inputs=2:duration=first[a]' \
-map 0:v -map '[a]' -c:v copy -c:a aac output.mp4How do I merge audio and video with the Sume API?
Render the clip over the new track with Timeline 1.0: the audio file becomes the render's spine, audio.url, and audio.duration_seconds sets the output length. In current code the render re-encodes the video, unlike -c:v copy, and leaves the clip's own sound out. Set output to the clip's size, since the default frame is 1080×1920. Replace or remove the audio in a video shows the request.
Where -shortest stops at the shorter stream, the render always follows the audio: the API reference says the output lasts exactly audio.duration_seconds.
- Video shorter than the audio: in current code, slots that end early hold their last frame to the end and warn
video_coverage_shorter_than_audio, and a slot longer than its source file is filled byrender.pad_mode, which loops short sources by default. Video freezes but audio continues covers both cases. - Video longer than the audio: end the last slot at the spine. In current code a slot that ends more than 0.5 seconds past it is refused with
invalid_segment_timing. - Both files must already be your workspace's
media.sume.comfiles, such as an audio detach WAV or a TTS master for the spine; which URLs each endpoint accepts explains the rule. In current code a source file over 300 MiB is refused withsource_too_large. - A render is listed at $0.10 per output minute on API pricing and reserves
ceil(audio.duration_seconds / 60)minutes, plus a 5.5% agent fee by default.
Sources
Related posts
More in Developers
- FFmpeg trim video: cut by time, with or without re-encoding
Trim a video with FFmpeg: -ss before -i to seek, -t for the length, and -c copy to skip re-encoding, which starts the cut at the keyframe before.
- Higgsfield API key: how to get one and send it
A Higgsfield API key is a key ID plus a secret made in Higgsfield Console, sent together in one Authorization: Key header from server code only.
- HMAC vs digital signature: what each proves for webhooks
An HMAC proves the sender holds a shared secret. A digital signature is made with a private key and checked with a public one, adding non-repudiation.
- How long does it take to generate an AI image?
On Sume, most AI images finish inside the 30 seconds the API holds a request open. What makes one take longer, and how to tell slow from stuck.
Written by Sume