Cut a clean voice reference from video: audio detach range, mono wav

Sume's audio detach takes a start and end range and can output mono wav at 16 kHz. Cut 10 seconds of a single speaker for $0.01, then check it before use.

5 min readSume
All posts

To cut a clean reference clip from a video, call Sume's audio detach with a range, format: "wav" and channels: "mono". It costs $0.01 per job, returns a durable audio_url, and takes up to 900 seconds of output per call. Pick ten to thirty seconds of one speaker with no music or crosstalk.

Why bother now: Microsoft says MAI-Voice-2.1 matches a voice from short reference clips with no fine-tuning (read 2026-10-03). Short references make clip quality the main variable, and the cheapest quality control is cutting the right seconds.

The request

The docs call 16000 with mono the speech to text shape. For a reference clip that a person will also audition, 44,100 or 48,000 keeps more of the voice. Which one a given voice engine prefers is something to check on that engine's own page; I did not read a requirement on the pages cited.

From Sume's Audio detach docs, read 2026-10-03.
FieldMeaning
video_urlThis workspace's media.sume.com artifact or asset; import first with POST /v1/media-imports
range{ start, end? } in seconds; end omitted means to the end
channelssource (default) or mono
sample_rate16000, 44100 or 48000; omit to inherit
formatwav or mp3
LimitsSource up to 1,800 s, output up to 900 s
Price$0.01 per job

Submit and read it back

The route is asynchronous by default. The snippet submits, polls the status route and prints the result. It sends a 12 second range starting at 41.5 seconds.

import os
import time
import uuid
import requests

BASE = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "range": {"start": 41.5, "end": 53.5},
    "format": "wav",
    "channels": "mono",
}
r = requests.post(f"{BASE}/v1/audio-detach", json=body, timeout=60,
                  headers={**H, "Idempotency-Key": f"ref-{uuid.uuid4()}"})
r.raise_for_status()
job_id = r.json().get("request_id") or r.json()["job"]["id"]
while True:
    s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=H, timeout=30).json()
    if s.get("result_ready") or s.get("status") in ("failed", "canceled"):
        break
    time.sleep(2)
print(requests.get(f"{BASE}/v1/jobs/{job_id}/result", headers=H, timeout=30).json())

Choosing the seconds

Listen for four things before you trust a range. No other voice, even faint. No music bed, including quiet room tone from a TV. No clipping. Natural pace, since the voice will inherit it. If the video has a music bed under the speech, Sume's docs list no stem splitter, so pick another stretch or ask for a clean recording.

Then, and before anything is cloned, confirm you have permission from the person speaking. Keep their written consent with the file. Microsoft's pages I read mention guardrails but not a consent mechanism, so the paper trail is yours to keep.

Checking the file you got

When the result arrives, read duration_seconds, channels and sample_rate from it and compare with what you asked for. A warnings[] array can come back too, so print it. If the duration is shorter than your range, the source ended first. If sample_rate is null, it was inherited from the source, which is expected when you omit it.

Then listen once on headphones. Thirty seconds of listening costs less than a cloned voice that carries a hum or a second speaker into every future line.

Costs and cleanup

The detach is a flat $0.01 whether the range is two seconds or 900. If you cut three candidate ranges and listen to all three, that is three cents. The file is a durable artf_ on media.sume.com, so you can reuse it without detaching again, but delete references you no longer have permission to use, and record where each one came from.

If the same speaker appears across a long video, a good habit is to cut one reference per setting: studio, outdoor, phone. Then audition the voice you get from each and keep the one that sounds closest in the setting where you will use it. A reference recorded in a quiet room often produces a cleaner voice than a longer one with background noise.

If you need several samples

For many ranges from one video, detach the whole track once and split it with timeline audio, which takes up to 20 ranges per job, rather than detaching again and again. Detach once, split once, and pay $0.02 instead of $0.01 times twenty.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume