Auphonic audio cleanup vs Sume audio detach: different jobs
Auphonic cleans and levels audio; Sume audio detach only pulls the track out of a video as wav or mp3. Compare billing, limits and where each belongs.

No: Sume does not clean up or level podcast audio the way Auphonic does. Sume audio detach extracts the audio track of one Sume-hosted video into a new wav or mp3 and leaves the video untouched; it runs worker ffmpeg with no processing options for noise or loudness.
Auphonic's pricing page (read 2026-10-10) describes noise reduction, intelligent leveling and filtering as its core. The two products sit at different points of a podcast workflow, and the sensible question is how they hand off, not which one wins. Sume facts here come from the audio detach docs.
What each product actually does
Auphonic describes itself as an audio processing platform with AI-based algorithms for podcast production: noise reduction, intelligent leveling, filtering, speech-to-text and video support. Sume audio detach is a demux step. It takes one media.sume.com video and returns a new audio artifact, sample-exact wav by default or 128 kbps mp3.
| Item | Auphonic | Sume audio detach |
|---|---|---|
| Core job | Noise reduction, leveling, filtering | Extract the audio track as a new file |
| Billing basis | Duration of audio sent, in milliseconds, minimum 3 minutes | $0.01 per job, no provider inference |
| Free allowance | 2 hours of processed audio per month | None stated in the docs |
| API | Full API, CLI, watch folders and batch on paid plans | POST /v1/audio-detach with a required Idempotency-Key |
| Length limits | Not stated on the pricing page | Source up to 1800 s, output up to 900 s |
| Output control | Multiple outputs at no extra charge | format wav or mp3, channels source or mono, sample_rate 16000, 44100 or 48000 |
Where the Sume step is useful
Detach earns its $0.01 when audio has to leave a video for another tool. The docs name three consumers: timeline_create takes it as audio.url, POST /v1/timeline-1.0/audio splits it, and speech-to-text reads it. The docs describe sample_rate: 16000 with channels: "mono" as the speech-to-text shape.
For a clip longer than 900 seconds you must pass a range, because the output cap is 900 s even though the source cap is 1800 s. A source with no audio fails with detach_source_has_no_audio, and the docs suggest checking probe.has_audio with video inspect first.
One rule surprises people who expect an ffmpeg passthrough: the server refuses af, filter, ffmpeg, cmd and codec fields with ffmpeg_fields_rejected. You cannot ask detach to normalize loudness, and the docs do not describe any such option.
How the two can sit in one workflow
A plausible order for a video podcast is: detach the audio from the recorded video, send the wav through Auphonic for leveling, then bring the result back into the edit. I have not verified that Sume accepts an externally produced audio file as a timeline spine; the docs say every URL must already be a workspace media.sume.com artifact and tell you to import files first with POST /v1/media-imports. Read Media inputs before you build around it.
If you only need the original audio, skip the second tool. Detach once, then use timeline audio to split the track into reusable pieces. The docs recommend detaching once for many ranges rather than detaching per range.
Cost shape
Auphonic's smallest billable unit is 3 minutes, so a 40 second voice note is billed as 3 minutes of its credit. The plans list 9, 21, 45, 100 and 250 hours per month for S through XXL, with a free tier of 2 hours. Sume's $0.01 is per job regardless of length inside the caps, so a 20-minute file and a 20-second file both cost one cent to detach.
Those numbers are not substitutes. One cent buys a file extraction, and Auphonic's hours buy processing. Compare them only when your pipeline needs both steps.
Sources
Related posts
More in Comparisons
- Canva Connect API vs the Sume API: what each is for
Canva Connect syncs designs, assets and comments, with some APIs in preview. The Sume API generates video, images and audio. Different jobs.
- CapCut lists Seedance 2.5 and Gemini Omni. Does Sume's API carry them?
CapCut's tools page names several video and image models. Sume's API documents Seedance 2.5 and Gemini Omni Flash 1.1; the rest are not in its docs.
- ChatGPT Image 2 vs 2.5 on Sume: $0.26375 vs $0.065875 and what differs
On Sume, GPT Image 2.5 high quality at 1024 costs $0.065875, a quarter of GPT Image 2 at $0.26375, and adds mask_url, background and 16 references.
- Creatomate RenderScript vs a Sume Timeline document: field map
Creatomate's RenderScript is a general scene JSON; Sume's Timeline 1.0 is one audio spine plus video slots. Field-by-field map and what Sume cannot express.
Written by Sume