Voice isolator vs audio detach: 500 MB and 1 hour vs 1800 seconds
ElevenLabs Voice Isolator removes background noise from speech. Sume audio detach extracts the video audio track as is; it does not separate vocals from music.

Voice isolation and audio detach are different jobs. ElevenLabs Voice Isolator removes background noise from a recording, while Sume audio detach copies the audio track out of a video into wav or mp3 without separating anything. If you need vocals separated from music, Sume does not do that today.
The confusion is common because both show up for the query "separate vocals from video". Here is what each does, with the limits each vendor states.
What ElevenLabs states
The Voice Isolator doc lists a 500 MB maximum and up to one hour of audio, accepts audio and video formats, and meters at 1000 characters per minute. It also says the tool is not specifically optimized for isolating vocals from music. On the API pricing page, Voice Isolator is listed at $0.12 per minute.
So even the vendor frames it as speech cleanup. If your goal is an acapella from a song, read that caveat first.
What Sume audio detach does
The audio detach docs describe POST /v1/audio-detach: give it a media.sume.com video and it returns that video's audio as wav (pcm_s16le by default) or mp3 at 128k. You can set a range with start and end, choose source or mono channels, and pick a sample rate of 16000, 44100 or 48000. The source can be up to 1800 seconds, the output up to 900 seconds, and it is $0.01 per job. A video with no audio returns detach_source_has_no_audio.
Nothing in that contract removes noise or splits stems. The output is the original soundtrack, cut to the range you asked for.
| Question | ElevenLabs Voice Isolator | Sume audio detach |
|---|---|---|
| What it does | Removes background noise from speech | Extracts the existing audio track |
| Input size | 500 MB, up to 1 hour | Source up to 1800 seconds |
| Output length | Matches input | Up to 900 seconds |
| Separates vocals from music | Not specifically optimized for it | No |
Pick by goal
Choose detach when you need the audio for something else: a transcript, a reference, a remix bed, or re-timing. The 16 kHz mono post shows the speech-to-text recipe. Choose a vendor isolator when the problem is noise on the voice track, and run that step outside Sume.
- Need a transcript: detach to 16 kHz mono, then transcribe.
- Need to mute the music under dialogue: use the timeline duck_db setting at render time instead of separating.
- Need a clean vocal stem: use a stem separation tool outside Sume.
- Need the original track for a re-cut: detach with a range.
Combine them
A practical chain is detach a range from the video, clean the voice with an external isolator, then bring the clean audio back into a timeline render as the audio url. Sume handles the cut and the final mix; the cleanup step lives where the tool does.
Sources
Related posts
More in Comparisons
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
- Wan 3.0 or Seedance 2.5: which is cheaper for a 10-second 720p clip?
Wan 3.0 is cheaper: a 10-second 720p clip is $1.00 at fal list ($1.25 on Sume) against about $4.62 ($5.78 on Sume) for Seedance 2.5, roughly 4.6x.
- Wan 3.0 vs Gemini Omni Flash at 720p: both list $0.10 per second
Wan 3.0 (fal) and Gemini Omni Flash (Google, about $0.10 per second at 720p) share a list price; Sume bills $0.125 per second for either.
- WellSaid cost per download minute: Pro annual $0.18, Starter $0.95
WellSaid Starter monthly is $19 for 20 download minutes ($0.95 each) and Pro annual is $33 for 180 ($0.18). Sume TTS is about $0.04 a minute at 900 characters.
Written by Sume