Strip the English voiceover before localizing: audio-detach is 1 cent

Pull the audio track from a finished ad with Sume audio-detach at $0.01 per job to reuse, review or replace it before a 12-market localization pass.

4 min readSume
All posts

Sume's audio detach takes one hosted video and returns its audio as a new artifact for $0.01 per job, so a 12-ad set is 12 x $0.01 = $0.12. It is the first step of a localization pass: get the English track out, send it to translation or review, then build the new language version on a timeline spine. Nothing in the job needs a model; the docs say only worker ffmpeg runs.

Required is video_url, which must be a media.sume.com artifact or asset in your workspace. The server does not fetch from the open internet, so import first.

What you get back

A finished job returns kind: audio_detach with audio_url (a new artifact), duration_seconds, format, channels and sample_rate. The default mode is async; pass mode: sync to wait up to 30 seconds for a completed job, otherwise you poll GET /v1/jobs/:id/status and /result.

audio-detach fields (read 2026-10-04, from the docs)
FieldValues
sample_rate16000, 44100, 48000, or omit to inherit
STT shape16000 with channels mono
Price$0.01 per job
Requiredvideo_url, Idempotency-Key

A localization pass in order

First detach the English track from each master. Second, send the audio or its transcript to a native reviewer. Third, make the new voice track for each market with your voice tool. Fourth, render each language on the same video slots by swapping only the spine in Timeline 1.0, $0.10 per ceil(output minute). For 12 markets of a 30 second ad, that is 12 x $0.10 = $1.20 of render cost.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: audio-detach-ad7-en-v1" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4"
  }'

Cost for the whole set

Add the steps for a 12-market, one-ad pass. The detach is $0.01 once, not per market. The timeline renders are $1.20. Voice generation and review are priced elsewhere, so this page does not total them.

Keep the detached English track in your records. It is the reference for a reviewer who needs to compare meaning, timing and legal text against each market version.

Fixed-rate steps for one ad into 12 markets (read 2026-10-04)
StepJobsRateCost
Detach English audio1$0.01$0.01
Timeline render per market, 30 s12$0.10$1.20
Subtotal of fixed-rate steps$1.21

If a market needs a longer script than the English, its render can pass 60 seconds and bill 2 minutes, which is $0.20, so check the length of each new track before you submit.

Mistakes to avoid

The most common mistake is detaching from a clip that has music and voice mixed. The job returns one audio artifact, so a reviewer hears both, and a translator may take the music for room noise. If you need the voice only, keep the original voice stem from production and use detach only for finished ads that have no stems left.

Another mistake is skipping the length check. A target language that is longer than English will not fit a 15 second slot unless the script is cut. Decide per market whether to shorten the script or lengthen the video, and make that call before you render.

A third is mixing up the keys. A key like audio-detach-ad7-en-v1 should name the ad, the language and a version. If you change the source video, bump the version, because the same key with a different body is refused. These small habits keep a 12-market holiday pass auditable.

Related posts

More in Media tools

All Media tools posts

Written by Sume