Cut hold music from a call recording before STT: one concat job

Drop a hold-music stretch with one Sume timeline audio concat job that reuses the file twice via source_in and duration, then send the result to STT.

5 min readSume
All posts

Short answer

Make one timeline audio concat job with two parts that point at the same recording: the part before the hold music, and the part after it. Each part takes source_in and duration, the result is one gapless file, and you send that file to STT so you do not pay to transcribe music. The job is $0.01 flat.

The fields are from the timeline audio page: parts[] takes 1 to 20 items, each { url, source_in?, duration? }, and the join is sample-domain with no silence at the seams.

Say a 12-minute call (720 s) has hold music from 3:10 to 6:40. STT takes at most 10 minutes per request, so removing the 210 seconds of music also brings the file under the limit: the parts 0 to 190 seconds and 400 to 720 seconds give 510 seconds. The numbers below are an invented example.

Worked example, invented timings
Partsource_in (s)duration (s)
Before hold0190
After hold400320
Joined length510 s (8.5 min)

The request body

The recording must already be audio on media.sume.com in your workspace; import it first with POST /v1/media-imports, which needs an Idempotency-Key.

{
  "operation": "concat",
  "parts": [
    { "url": "https://media.sume.com/artifacts/artf_demo/call.wav", "source_in": 0, "duration": 190 },
    { "url": "https://media.sume.com/artifacts/artf_demo/call.wav", "source_in": 400, "duration": 320 }
  ],
  "output": { "format": "wav" }
}

Mapping word times back

The result carries segments[] with the offset of each part in the new file. Keep them: a word that STT places at 200 seconds in the joined file sits at 410 seconds in the original, because it falls in the second part and 210 seconds were removed. Part two starts at 190 in the new file and at 400 in the original, a difference of 210 seconds.

Format and errors

Keep the format on wav. The docs say mp3 adds priming padding at every edge, which matters at a seam. Concat parts must share a channel layout, and the same file used twice does.

Cost

Total cost is the $0.01 concat job plus STT at $0.01 per audio minute on the joined file, so the 8.5 minutes above come to about 9 cents; the 3.5 minutes of music you skipped would have cost about 4 cents more, and the original would not fit in one request if it were longer than 10 minutes.

Finding the cut points

You need the start and end of the hold stretch. If you have a transcript, the gap in words[] is the clue: a run of seconds with no words, or with garbled ones, between two phone-line phrases. Otherwise listen once and note the times, since the cuts are only as accurate as your times.

Give yourself a small margin. Cut a few tenths of a second inside the music at both ends rather than inside speech, because a clipped first word is worse than half a second of music left in.

If a call has more than one hold, add more parts to the same job; a concat job takes up to 20. Each extra part is just another source_in and duration pair on the same URL, and the price stays $0.01 for the job.

  • Take cut times from the word gaps, or by listening.
  • Cut inside the music, not inside speech.
  • Several holds mean several parts, up to 20.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume