Cut a clean voice reference from video: audio detach range, mono wav
Sume's audio detach takes a start and end range and can output mono wav at 16 kHz. Cut 10 seconds of a single speaker for $0.01, then check it before use.

To cut a clean reference clip from a video, call Sume's audio detach with a range, format: "wav" and channels: "mono". It costs $0.01 per job, returns a durable audio_url, and takes up to 900 seconds of output per call. Pick ten to thirty seconds of one speaker with no music or crosstalk.
Why bother now: Microsoft says MAI-Voice-2.1 matches a voice from short reference clips with no fine-tuning (read 2026-10-03). Short references make clip quality the main variable, and the cheapest quality control is cutting the right seconds.
The request
The docs call 16000 with mono the speech to text shape. For a reference clip that a person will also audition, 44,100 or 48,000 keeps more of the voice. Which one a given voice engine prefers is something to check on that engine's own page; I did not read a requirement on the pages cited.
| Field | Meaning |
|---|---|
video_url | This workspace's media.sume.com artifact or asset; import first with POST /v1/media-imports |
range | { start, end? } in seconds; end omitted means to the end |
channels | source (default) or mono |
sample_rate | 16000, 44100 or 48000; omit to inherit |
format | wav or mp3 |
| Limits | Source up to 1,800 s, output up to 900 s |
| Price | $0.01 per job |
Submit and read it back
The route is asynchronous by default. The snippet submits, polls the status route and prints the result. It sends a 12 second range starting at 41.5 seconds.
import os
import time
import uuid
import requests
BASE = "https://api.sume.com"
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
body = {
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"range": {"start": 41.5, "end": 53.5},
"format": "wav",
"channels": "mono",
}
r = requests.post(f"{BASE}/v1/audio-detach", json=body, timeout=60,
headers={**H, "Idempotency-Key": f"ref-{uuid.uuid4()}"})
r.raise_for_status()
job_id = r.json().get("request_id") or r.json()["job"]["id"]
while True:
s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=H, timeout=30).json()
if s.get("result_ready") or s.get("status") in ("failed", "canceled"):
break
time.sleep(2)
print(requests.get(f"{BASE}/v1/jobs/{job_id}/result", headers=H, timeout=30).json())
Choosing the seconds
Listen for four things before you trust a range. No other voice, even faint. No music bed, including quiet room tone from a TV. No clipping. Natural pace, since the voice will inherit it. If the video has a music bed under the speech, Sume's docs list no stem splitter, so pick another stretch or ask for a clean recording.
Then, and before anything is cloned, confirm you have permission from the person speaking. Keep their written consent with the file. Microsoft's pages I read mention guardrails but not a consent mechanism, so the paper trail is yours to keep.
Checking the file you got
When the result arrives, read duration_seconds, channels and sample_rate from it and compare with what you asked for. A warnings[] array can come back too, so print it. If the duration is shorter than your range, the source ended first. If sample_rate is null, it was inherited from the source, which is expected when you omit it.
Then listen once on headphones. Thirty seconds of listening costs less than a cloned voice that carries a hum or a second speaker into every future line.
Costs and cleanup
The detach is a flat $0.01 whether the range is two seconds or 900. If you cut three candidate ranges and listen to all three, that is three cents. The file is a durable artf_ on media.sume.com, so you can reuse it without detaching again, but delete references you no longer have permission to use, and record where each one came from.
If the same speaker appears across a long video, a good habit is to cut one reference per setting: studio, outdoor, phone. Then audition the voice you get from each and keep the one that sounds closest in the setting where you will use it. A reference recorded in a quiet room often produces a cleaner voice than a longer one with background noise.
If you need several samples
For many ranges from one video, detach the whole track once and split it with timeline audio, which takes up to 20 ranges per job, rather than detaching again and again. Detach once, split once, and pay $0.02 instead of $0.01 times twenty.
Sources
Related posts
More in Media tools
- Cyber Monday half-banner: price card over a clip, compose stack
Stack a Sume-hosted price card above a product clip in one vertical shot with Timeline compose, at a flat $0.02 per job, then drop it into a Timeline.
- Demand Gen video under 10 seconds: no in-stream; upload 1 to 5 videos
Demand Gen needs 5 seconds minimum, but videos under 10 seconds can't serve on YouTube in-stream. Upload 1 to 5 per ad; crop a 9:16 master on Sume.
- Demucs is archived: still use it to split a song into stems?
Demucs was archived on January 1, 2025 but is MIT-licensed and still splits drums, bass, vocals and other. How to feed it a track from Sume audio detach.
- Descript AI music and sound effects vs generating music by API
Descript's 2026-09-17 update generates music and sound effects on request. Sume's Music Router generates music from a prompt; it has no sound-effect route.
Written by Sume