Mixed-language recording: detach ranges, STT with a hint per range
Split a bilingual video into audio ranges with Sume audio detach, then transcribe each range with its own language_code. Detach output caps at 900 s.

For a recording that switches languages, cut it into ranges, then transcribe each range with the right language_code. On Sume, detach the audio once, then split it with timeline audio split, then run speech-to-text on each piece with its own hint. Microsoft's MAI-Transcribe-2-Streaming is reported to offer continuous language detection across 60 languages (tracker, not the vendor, read 2026-10-05). Sume's STT takes a hint per request, so a mixed recording is a pipeline.
The pipeline
| Step | Call | Notes |
|---|---|---|
| 1. Detach once | POST /v1/audio-detach | wav, mono, 16000 is the STT shape; $0.01 |
| 2. Split | POST /v1/timeline-1.0/audio, operation split | 1 to 20 ranges, each start and optional end; $0.01 |
| 3. Transcribe each piece | STT with language_code | Hint per piece; auto-detect if omitted |
| 4. Merge | Your code | Add each piece's start offset to its word times |
Details
You need to know where the language changes. Either ask the speaker, or transcribe once with auto-detect, read the language_code and word times, then re-cut at the boundaries.
- Detach output is at most 900 s, and the source at most 1800 s. For a longer track, use a
range. - Split ranges can overlap, and an omitted
endruns to the end of the file. - Without
duration_seconds, STT reserves 1 minute. Pass each piece's real length. - Timeline audio has no resource GET. Read
/v1/jobs/:id/result.
When one pass is enough
If the second language is a few words, an en hint with script_text captions is simpler than splitting. Split when whole stretches change language, because a wrong hint on a long range is the expensive mistake.
Cost shape
Count the jobs before you start. One detach at $0.01, one split at $0.01, and STT at $0.01 per audio minute across the pieces. Ten minutes of audio is about ten cents of STT plus two cents of ffmpeg jobs, using the public rates in the docs. Read GET /v1/catalog for the live rates.
Detach once even if you need many pieces. The audio detach page says so directly: for many ranges, detach once and then split with timeline audio. Repeating detach for each range pays for the same demux again.
Merge the transcripts in your own code. Each piece starts at zero, so add its start offset to every word and segment time, then send the merged words to POST /v1/video-captions. Sume will burn them without a second speech-to-text pass.
Sources
Related posts
More in Use cases
- MSN's AI content policy: AIAC vs unreviewed AIGC for video makers
MSN separates AI-assisted content from unreviewed AI-generated content, and asks for disclosure now. What a Sume video publisher on MSN should do.
- A holiday music video from one Lyria track: 18 clips, about $23.50
Lyria track $0.125, eighteen 10-second Omni Flash clips at 720p $22.50, Timeline $0.30, captions $0.60. Total about $23.53 for 3 minutes.
- Music visualizer clip: MiniMax H3 Max, image plus audio reference
Make a 10-second visualizer from cover art plus a track excerpt with minimax-h3-max on Sume: $1.00 at 768p; audio alone is not a valid reference.
- Napkin sketch to product render: one reference on GPT Image 2.5
Turn a photographed napkin sketch into a product render with one input reference on GPT Image 2.5. Set aspect_ratio auto to keep the shape. Code and prices.
Written by Sume