Is this video stereo or mono? Read probe audio_channels first
Sume video inspect returns audio_channels, audio_sample_rate and audio_codec. Use them to choose channels and sample_rate for audio detach before you submit.

To find out whether a video carries stereo or mono sound, run Sume video inspect and read probe.audio_channels (an integer, or null), plus audio_sample_rate and audio_codec. These come from the same ffprobe pass as the video fields, and has_audio tells you whether there is a track at all.
Check rather than assume. A mono dialogue clip and a two-channel music-bed clip both come out of one detach call, but they want different settings.
Choose detach settings from the probe
| Goal | channels | sample_rate | format |
|---|---|---|---|
| Speech to text | mono | 16000 | wav |
| Keep a stereo mix | source (default) | omit, inherits source | wav |
| Small file to share | source | 44100 or 48000 | mp3 (128 kbps) |
Why the defaults matter
Audio detach returns sample-exact wav (pcm_s16le) by default, and channels defaults to source, so a stereo track stays stereo. Setting channels: "mono" downmixes. sample_rate takes only 16000, 44100 or 48000, and 16000 with mono is the speech-to-text shape.
The output cap is 900 seconds, and the source cap is 1800 seconds. A whole track longer than 900 seconds needs a range. Public rate is $0.01 per job.
Two lines of logic
Read the probe, then decide.
probe = result["probe"]
if not probe["has_audio"]:
raise SystemExit("no audio track: skip detach")
body = {"video_url": url}
if probe["audio_channels"] and probe["audio_channels"] > 1:
# keep stereo for a mix; downmix only for speech-to-text
body["channels"] = "source"
print(body)Detach once, then split many ranges later with timeline audio, rather than detaching the same video repeatedly.
Sources
Related posts
More in Media tools
- Keep a whole 16:9 shot in 9:16: blur fill instead of cropping
Cropping loses 68% of a 16:9 frame. Timeline fit blur keeps the full shot centered over a blurred copy. How it works on Sume and when to prefer it.
- Kling Motion Control character_orientation: video or image?
Pick 'video' (the default) for complex motion up to 30 s; 'image' follows camera movement but fal documents a 10 s limit. Sume bills $0.1575 per second.
- Kling Motion Control prompt: appearance only, motion is the video
On Sume's Kling 3.0 Motion Control the prompt guides appearance only. Motion comes from motion_video_url, and the body rejects reference_image_urls and model.
- Lip-sync a saved avatar with H3 Max: avatar_handle, not image_url
Use a ready Sume avatar as the face for MiniMax H3 Max Lip Sync: send avatar_handle plus Sume-hosted audio, 5-14.8 seconds, and see the price at 768p.
Written by Sume