Auphonic audio cleanup vs Sume audio detach: different jobs

Auphonic cleans and levels audio; Sume audio detach only pulls the track out of a video as wav or mp3. Compare billing, limits and where each belongs.

5 min readSume
All posts

No: Sume does not clean up or level podcast audio the way Auphonic does. Sume audio detach extracts the audio track of one Sume-hosted video into a new wav or mp3 and leaves the video untouched; it runs worker ffmpeg with no processing options for noise or loudness.

Auphonic's pricing page (read 2026-10-10) describes noise reduction, intelligent leveling and filtering as its core. The two products sit at different points of a podcast workflow, and the sensible question is how they hand off, not which one wins. Sume facts here come from the audio detach docs.

What each product actually does

Auphonic describes itself as an audio processing platform with AI-based algorithms for podcast production: noise reduction, intelligent leveling, filtering, speech-to-text and video support. Sume audio detach is a demux step. It takes one media.sume.com video and returns a new audio artifact, sample-exact wav by default or 128 kbps mp3.

Auphonic from its pricing page (read 2026-10-10); Sume from the audio detach docs
ItemAuphonicSume audio detach
Core jobNoise reduction, leveling, filteringExtract the audio track as a new file
Billing basisDuration of audio sent, in milliseconds, minimum 3 minutes$0.01 per job, no provider inference
Free allowance2 hours of processed audio per monthNone stated in the docs
APIFull API, CLI, watch folders and batch on paid plansPOST /v1/audio-detach with a required Idempotency-Key
Length limitsNot stated on the pricing pageSource up to 1800 s, output up to 900 s
Output controlMultiple outputs at no extra chargeformat wav or mp3, channels source or mono, sample_rate 16000, 44100 or 48000

Where the Sume step is useful

Detach earns its $0.01 when audio has to leave a video for another tool. The docs name three consumers: timeline_create takes it as audio.url, POST /v1/timeline-1.0/audio splits it, and speech-to-text reads it. The docs describe sample_rate: 16000 with channels: "mono" as the speech-to-text shape.

For a clip longer than 900 seconds you must pass a range, because the output cap is 900 s even though the source cap is 1800 s. A source with no audio fails with detach_source_has_no_audio, and the docs suggest checking probe.has_audio with video inspect first.

One rule surprises people who expect an ffmpeg passthrough: the server refuses af, filter, ffmpeg, cmd and codec fields with ffmpeg_fields_rejected. You cannot ask detach to normalize loudness, and the docs do not describe any such option.

How the two can sit in one workflow

A plausible order for a video podcast is: detach the audio from the recorded video, send the wav through Auphonic for leveling, then bring the result back into the edit. I have not verified that Sume accepts an externally produced audio file as a timeline spine; the docs say every URL must already be a workspace media.sume.com artifact and tell you to import files first with POST /v1/media-imports. Read Media inputs before you build around it.

If you only need the original audio, skip the second tool. Detach once, then use timeline audio to split the track into reusable pieces. The docs recommend detaching once for many ranges rather than detaching per range.

Cost shape

Auphonic's smallest billable unit is 3 minutes, so a 40 second voice note is billed as 3 minutes of its credit. The plans list 9, 21, 45, 100 and 250 hours per month for S through XXL, with a free tier of 2 hours. Sume's $0.01 is per job regardless of length inside the caps, so a 20-minute file and a 20-second file both cost one cent to detach.

Those numbers are not substitutes. One cent buys a file extraction, and Auphonic's hours buy processing. Compare them only when your pipeline needs both steps.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume