FFmpeg concatenate videos: concat demuxer vs concat filter

Concatenate videos with FFmpeg: the concat demuxer joins matching files with -c copy and no re-encode; the concat filter re-encodes mixed clips.

6 min readSume
All posts

To concatenate videos with FFmpeg, use the concat demuxer when every file has the same streams and settings: list the files in a text file and run ffmpeg -f concat -safe 0 -i list.txt -c copy output.mp4, which joins them without re-encoding. When the clips differ in size, frame rate, or codec, use the concat filter instead: it decodes the clips, joins them, and encodes the result.

The FFmpeg facts come from its formats, filters, and ffmpeg documentation, and the Sume facts from the Timeline 1.0 and Audio detach docs, all read on 2026-09-28. Anything described as current behavior is read from Sume's code.

How do I concatenate MP4 files without re-encoding?

Write a list file with one file line per clip, in playback order, then read it with the concat demuxer and copy the streams:

  • The concat demuxer reads the files one after the other, as if all their packets had been muxed together. Timestamps are shifted so each file starts where the previous one finishes.
  • -f concat selects the demuxer. A first line of ffconcat version 1.0 lets FFmpeg recognize the script on its own instead.
  • -safe 0 accepts any file name. By default FFmpeg rejects paths that are absolute, name a protocol, or use characters outside letters, digits, period, underscore, and hyphen.
  • All files must have the same streams: same codecs, same time base, and so on. The shift is global, so streams of unequal length inside a file can leave gaps, and a wrong file duration can cause artifacts.
  • -c copy means no decoding or encoding, so there is no quality loss.
# list.txt
file 'part1.mp4'
file 'part2.mp4'
file 'part3.mp4'

ffmpeg -f concat -safe 0 -i list.txt -c copy output.mp4

How do I concatenate videos with different resolutions or codecs?

Use the concat filter, which works on decoded frames, so each input can be scaled to one size first. This joins two clips with sound into a 1920×1080 file, fitting each inside the frame with black bars where its shape differs:

  • concat=n=2:v=1:a=1 joins two segments of one video and one audio stream each. Every segment must have the same number of streams of each type, so a clip without sound breaks this join.
  • FFmpeg picks a common pixel format, sample rate, and channel layout on its own, but other settings, such as resolution, must be converted explicitly. That is the scale and pad step.
  • Different frame rates are accepted but give variable frame rate output, so the example sets one rate with the fps filter.
  • The filter uses the longest stream in each segment except the last, and pads shorter audio with silence.
  • The output is re-encoded, and FFmpeg's docs say encoding in most cases degrades quality.
ffmpeg -i a.mp4 -i b.mp4 -filter_complex \
'[0:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v0];
 [1:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v1];
 [v0][0:a][v1][1:a]concat=n=2:v=1:a=1[v][a]' \
-map '[v]' -map '[a]' -c:v libx264 -c:a aac output.mp4

Which method should I use?

Copy when every file has the same streams, for example parts cut from one recording with a stream copy. Re-encode when they don't: the demuxer requires matching codecs and time bases, and changing a clip's size or rate takes a filter, which a stream copy can't apply.

FFmpeg behavior from the formats and filters documentation; Sume from Timeline 1.0 and its compiler as it runs today, read 2026-09-28.
MethodRe-encodesClips must matchSound
Concat demuxer with -c copyNoYes: same streams, codecs, and time baseEach file's own
Concat filterYesOnly in stream counts; you convert sizes firstEach segment's own; shorter audio padded with silence
Sume Timeline 1.0 renderYes, with libx264 in current codeNo: each slot is fitted to one output sizeOnly the render's audio spine and soundtrack

Can I concatenate videos with the Sume API?

Yes, but only the re-encoding way. A Timeline 1.0 render places 1 to 200 video[] slots back to back in one MP4 of up to 1,800 seconds and fits each clip to one output size with fit (cover, contain, stretch, or blur). In current code it always re-encodes the video with libx264; there is no stream-copy mode. How to merge two videos of different resolutions shows the full request.

  • In current code the render's sound comes only from its audio spine and an optional soundtrack, never from the clips. To keep each clip's sound, detach it with POST /v1/audio-detach and pass the files in order as audio.parts[], up to 20.
  • Every URL must already be your workspace's media.sume.com artifact or asset, such as an earlier Sume job's output; which URLs each endpoint accepts explains the rule. In current code a source file over 300 MiB is refused with source_too_large, and a render whose files add up to more than 4 GiB with download_budget_exceeded.
  • A render is listed at $0.10 per output minute on API pricing and reserves ceil(audio.duration_seconds / 60) minutes, plus a 5.5% agent fee by default. For long joins, see assembling a long-form video.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume