Voice isolator vs audio detach: 500 MB and 1 hour vs 1800 seconds

ElevenLabs Voice Isolator removes background noise from speech. Sume audio detach extracts the video audio track as is; it does not separate vocals from music.

4 min readSume
All posts

Voice isolation and audio detach are different jobs. ElevenLabs Voice Isolator removes background noise from a recording, while Sume audio detach copies the audio track out of a video into wav or mp3 without separating anything. If you need vocals separated from music, Sume does not do that today.

The confusion is common because both show up for the query "separate vocals from video". Here is what each does, with the limits each vendor states.

What ElevenLabs states

The Voice Isolator doc lists a 500 MB maximum and up to one hour of audio, accepts audio and video formats, and meters at 1000 characters per minute. It also says the tool is not specifically optimized for isolating vocals from music. On the API pricing page, Voice Isolator is listed at $0.12 per minute.

So even the vendor frames it as speech cleanup. If your goal is an acapella from a song, read that caveat first.

What Sume audio detach does

The audio detach docs describe POST /v1/audio-detach: give it a media.sume.com video and it returns that video's audio as wav (pcm_s16le by default) or mp3 at 128k. You can set a range with start and end, choose source or mono channels, and pick a sample rate of 16000, 44100 or 48000. The source can be up to 1800 seconds, the output up to 900 seconds, and it is $0.01 per job. A video with no audio returns detach_source_has_no_audio.

Nothing in that contract removes noise or splits stems. The output is the original soundtrack, cut to the range you asked for.

Isolation versus detach, side by side (read 2026-10-03)
QuestionElevenLabs Voice IsolatorSume audio detach
What it doesRemoves background noise from speechExtracts the existing audio track
Input size500 MB, up to 1 hourSource up to 1800 seconds
Output lengthMatches inputUp to 900 seconds
Separates vocals from musicNot specifically optimized for itNo

Pick by goal

Choose detach when you need the audio for something else: a transcript, a reference, a remix bed, or re-timing. The 16 kHz mono post shows the speech-to-text recipe. Choose a vendor isolator when the problem is noise on the voice track, and run that step outside Sume.

  • Need a transcript: detach to 16 kHz mono, then transcribe.
  • Need to mute the music under dialogue: use the timeline duck_db setting at render time instead of separating.
  • Need a clean vocal stem: use a stem separation tool outside Sume.
  • Need the original track for a re-cut: detach with a range.

Combine them

A practical chain is detach a range from the video, clean the voice with an external isolator, then bring the clean audio back into a timeline render as the audio url. Sume handles the cut and the final mix; the cleanup step lives where the tool does.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume