FFmpeg extract audio from video: copy, MP3, or WAV
Extract audio from a video with FFmpeg: -vn drops the picture, then -c:a copy keeps the audio as it is, or an encoder writes WAV or MP3.

To extract audio from a video with FFmpeg, drop the picture with -vn and choose what happens to the sound: copy the audio stream as it is with -c:a copy, which keeps its codec with no re-encode and no quality loss, or encode it to a new format such as WAV (-c:a pcm_s16le) or MP3 (-c:a libmp3lame).
The FFmpeg facts come from its ffmpeg, codecs, formats, and utilities documentation, and the Sume facts from the Audio detach and Video inspect docs, all read on 2026-09-28. Anything described as current behavior is read from Sume's code.
How do I extract audio without re-encoding it?
Copy the audio stream into an audio-only file. FFmpeg lists m4a among the MP4 family's file extensions, so AAC audio from an MP4 can go into .m4a:
-vn, as an output option, disables video recording: no video stream is selected or mapped for the output.copytells FFmpeg the stream is not to be re-encoded. With no decoding or encoding, FFmpeg's docs say, there is no quality loss.- A copy can fail. FFmpeg's docs say stream copy might not work in some cases, for example when information the target container needs isn't in the source, so the output extension has to be a container that can hold the codec you copy.
- Cutting while copying is loose: FFmpeg's docs say most formats can't seek exactly, so with
-ssbefore-iit seeks to the closest seek point before the position, and a stream copy keeps that extra segment instead of discarding it.
ffmpeg -i input.mp4 -vn -c:a copy output.m4aHow do I extract audio to WAV or MP3?
Name an audio encoder instead of copy. Encoding decodes the audio and encodes it again, which FFmpeg's docs say in most cases degrades quality, so copy when the original codec will do:
libmp3lamewraps the LAME MP3 encoder, and FFmpeg has it only when built with--enable-libmp3lame.-b:a 128kis the bitrate FFmpeg's own MP3 example uses.-ac 1sets one channel (mono) and-ar 16000sets the sample rate. Without them, FFmpeg keeps the input's channel count and sample rate.- To stay in AAC,
-c:a aacis FFmpeg's native AAC encoder, which uses 128 kbps when you don't set a bitrate.
# WAV: uncompressed PCM, signed 16-bit little-endian
ffmpeg -i input.mp4 -vn -c:a pcm_s16le output.wav
# MP3 at 128 kbps
ffmpeg -i input.mp4 -vn -c:a libmp3lame -b:a 128k output.mp3How do I pick one audio track or a time range?
-map 0:a:1takes the second audio stream of the first input; counting starts at 0, so FFmpeg's docs select the third audio stream with0:a:2.-ss 30before-istarts at 30 seconds, and-t 10keeps 10 seconds. Times can be plain seconds such as55or0.2, orHH:MM:SS.- When you encode, the start is exact: with
-accurate_seek, the default, FFmpeg decodes and discards the part between the seek point and the position.
Can I extract audio with the Sume API instead?
Yes, for a video already in your workspace: POST /v1/audio-detach runs the encode path of these commands on the server. It never copies the source stream; it always writes a new WAV or MP3 file. In current code the options map to FFmpeg like this:
| To | FFmpeg | Sume audio detach |
|---|---|---|
| Drop the picture | -vn | Always, in current code |
| Take the first audio stream | -map 0:a:0 | Always, in current code |
| Write 16-bit WAV | -c:a pcm_s16le | format: "wav", the default |
| Write MP3 | -c:a libmp3lame -b:a 128k | format: "mp3", at 128 kbps |
| Mix to mono | -ac 1 | channels: "mono" |
| Set the sample rate | -ar 16000 | sample_rate: 16000, 44100, or 48000 |
| Keep a time range | -ss 30 -t 10 | range with start: 30 and end: 40 |
| Copy without re-encoding | -c:a copy | Not offered |
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: audio-detach-mp3-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"format": "mp3",
"range": { "start": 30, "end": 40 }
}'What are the limits of audio detach?
video_urlmust be your workspace'smedia.sume.comartifact or asset, such as an earlier Sume job's output; which URLs each endpoint accepts explains the rule.- Sources run up to 1,800 seconds and each output up to 900 seconds, so a whole track past 900 seconds needs a
range. In current code a source file over 300 MiB is refused withsource_too_large. - A video with no audio track fails with
detach_source_has_no_audio. Checkprobe.has_audiofirst with aframes: falsevideo inspect. - Detach is billed per job, plus a 5.5% agent fee by default; the docs say to confirm the rate in
GET /v1/catalog. Trim, filter, or detach audio covers the rest of the call.
Sources
Related posts
More in Developers
- FFmpeg extract frames from video: one, every second, or all
Extract frames with FFmpeg: -frames:v 1 saves one image, -r 1 saves one per second, and a numbered file pattern writes every frame.
- FFmpeg height not divisible by 2: why and how to fix it
The error means a yuv420p encode got an odd height (or width). Make both even: scale with -2 or round with trunc(ih/2)*2. How Sume rounds sizes.
- FFmpeg trim video: cut by time, with or without re-encoding
Trim a video with FFmpeg: -ss before -i to seek, -t for the length, and -c copy to skip re-encoding, which starts the cut at the keyframe before.
- Higgsfield API key: how to get one and send it
A Higgsfield API key is a key ID plus a secret made in Higgsfield Console, sent together in one Authorization: Key header from server code only.
Written by Sume