FFmpeg concatenate videos: concat demuxer vs concat filter
Concatenate videos with FFmpeg: the concat demuxer joins matching files with -c copy and no re-encode; the concat filter re-encodes mixed clips.

To concatenate videos with FFmpeg, use the concat demuxer when every file has the same streams and settings: list the files in a text file and run ffmpeg -f concat -safe 0 -i list.txt -c copy output.mp4, which joins them without re-encoding. When the clips differ in size, frame rate, or codec, use the concat filter instead: it decodes the clips, joins them, and encodes the result.
The FFmpeg facts come from its formats, filters, and ffmpeg documentation, and the Sume facts from the Timeline 1.0 and Audio detach docs, all read on 2026-09-28. Anything described as current behavior is read from Sume's code.
How do I concatenate MP4 files without re-encoding?
Write a list file with one file line per clip, in playback order, then read it with the concat demuxer and copy the streams:
- The concat demuxer reads the files one after the other, as if all their packets had been muxed together. Timestamps are shifted so each file starts where the previous one finishes.
-f concatselects the demuxer. A first line offfconcat version 1.0lets FFmpeg recognize the script on its own instead.-safe 0accepts any file name. By default FFmpeg rejects paths that are absolute, name a protocol, or use characters outside letters, digits, period, underscore, and hyphen.- All files must have the same streams: same codecs, same time base, and so on. The shift is global, so streams of unequal length inside a file can leave gaps, and a wrong file duration can cause artifacts.
-c copymeans no decoding or encoding, so there is no quality loss.
# list.txt
file 'part1.mp4'
file 'part2.mp4'
file 'part3.mp4'
ffmpeg -f concat -safe 0 -i list.txt -c copy output.mp4How do I concatenate videos with different resolutions or codecs?
Use the concat filter, which works on decoded frames, so each input can be scaled to one size first. This joins two clips with sound into a 1920×1080 file, fitting each inside the frame with black bars where its shape differs:
concat=n=2:v=1:a=1joins two segments of one video and one audio stream each. Every segment must have the same number of streams of each type, so a clip without sound breaks this join.- FFmpeg picks a common pixel format, sample rate, and channel layout on its own, but other settings, such as resolution, must be converted explicitly. That is the
scaleandpadstep. - Different frame rates are accepted but give variable frame rate output, so the example sets one rate with the
fpsfilter. - The filter uses the longest stream in each segment except the last, and pads shorter audio with silence.
- The output is re-encoded, and FFmpeg's docs say encoding in most cases degrades quality.
ffmpeg -i a.mp4 -i b.mp4 -filter_complex \
'[0:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v0];
[1:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v1];
[v0][0:a][v1][1:a]concat=n=2:v=1:a=1[v][a]' \
-map '[v]' -map '[a]' -c:v libx264 -c:a aac output.mp4Which method should I use?
Copy when every file has the same streams, for example parts cut from one recording with a stream copy. Re-encode when they don't: the demuxer requires matching codecs and time bases, and changing a clip's size or rate takes a filter, which a stream copy can't apply.
| Method | Re-encodes | Clips must match | Sound |
|---|---|---|---|
Concat demuxer with -c copy | No | Yes: same streams, codecs, and time base | Each file's own |
| Concat filter | Yes | Only in stream counts; you convert sizes first | Each segment's own; shorter audio padded with silence |
| Sume Timeline 1.0 render | Yes, with libx264 in current code | No: each slot is fitted to one output size | Only the render's audio spine and soundtrack |
Can I concatenate videos with the Sume API?
Yes, but only the re-encoding way. A Timeline 1.0 render places 1 to 200 video[] slots back to back in one MP4 of up to 1,800 seconds and fits each clip to one output size with fit (cover, contain, stretch, or blur). In current code it always re-encodes the video with libx264; there is no stream-copy mode. How to merge two videos of different resolutions shows the full request.
- In current code the render's sound comes only from its audio spine and an optional soundtrack, never from the clips. To keep each clip's sound, detach it with
POST /v1/audio-detachand pass the files in order asaudio.parts[], up to 20. - Every URL must already be your workspace's
media.sume.comartifact or asset, such as an earlier Sume job's output; which URLs each endpoint accepts explains the rule. In current code a source file over 300 MiB is refused withsource_too_large, and a render whose files add up to more than 4 GiB withdownload_budget_exceeded. - A render is listed at $0.10 per output minute on API pricing and reserves
ceil(audio.duration_seconds / 60)minutes, plus a 5.5% agent fee by default. For long joins, see assembling a long-form video.
Sources
Related posts
More in Developers
- FFmpeg extract audio from video: copy, MP3, or WAV
Extract audio from a video with FFmpeg: -vn drops the picture, then -c:a copy keeps the audio as it is, or an encoder writes WAV or MP3.
- FFmpeg extract frames from video: one, every second, or all
Extract frames with FFmpeg: -frames:v 1 saves one image, -r 1 saves one per second, and a numbered file pattern writes every frame.
- FFmpeg height not divisible by 2: why and how to fix it
The error means a yuv420p encode got an odd height (or width). Make both even: scale with -2 or round with trunc(ih/2)*2. How Sume rounds sizes.
- Higgsfield API key: how to get one and send it
A Higgsfield API key is a key ID plus a secret made in Higgsfield Console, sent together in one Authorization: Key header from server code only.
Written by Sume