Mixed-language recording: detach ranges, STT with a hint per range

Split a bilingual video into audio ranges with Sume audio detach, then transcribe each range with its own language_code. Detach output caps at 900 s.

5 min readSume
All posts

For a recording that switches languages, cut it into ranges, then transcribe each range with the right language_code. On Sume, detach the audio once, then split it with timeline audio split, then run speech-to-text on each piece with its own hint. Microsoft's MAI-Transcribe-2-Streaming is reported to offer continuous language detection across 60 languages (tracker, not the vendor, read 2026-10-05). Sume's STT takes a hint per request, so a mixed recording is a pipeline.

The pipeline

Mixed-language audio pipeline on Sume (read 2026-10-05)
StepCallNotes
1. Detach oncePOST /v1/audio-detachwav, mono, 16000 is the STT shape; $0.01
2. SplitPOST /v1/timeline-1.0/audio, operation split1 to 20 ranges, each start and optional end; $0.01
3. Transcribe each pieceSTT with language_codeHint per piece; auto-detect if omitted
4. MergeYour codeAdd each piece's start offset to its word times

Details

You need to know where the language changes. Either ask the speaker, or transcribe once with auto-detect, read the language_code and word times, then re-cut at the boundaries.

  • Detach output is at most 900 s, and the source at most 1800 s. For a longer track, use a range.
  • Split ranges can overlap, and an omitted end runs to the end of the file.
  • Without duration_seconds, STT reserves 1 minute. Pass each piece's real length.
  • Timeline audio has no resource GET. Read /v1/jobs/:id/result.

When one pass is enough

If the second language is a few words, an en hint with script_text captions is simpler than splitting. Split when whole stretches change language, because a wrong hint on a long range is the expensive mistake.

Cost shape

Count the jobs before you start. One detach at $0.01, one split at $0.01, and STT at $0.01 per audio minute across the pieces. Ten minutes of audio is about ten cents of STT plus two cents of ffmpeg jobs, using the public rates in the docs. Read GET /v1/catalog for the live rates.

Detach once even if you need many pieces. The audio detach page says so directly: for many ranges, detach once and then split with timeline audio. Repeating detach for each range pays for the same demux again.

Merge the transcripts in your own code. Each piece starts at zero, so add its start offset to every word and segment time, then send the merged words to POST /v1/video-captions. Sume will burn them without a second speech-to-text pass.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume