Cut hold music from a call recording before STT: one concat job
Drop a hold-music stretch with one Sume timeline audio concat job that reuses the file twice via source_in and duration, then send the result to STT.

Short answer
Make one timeline audio concat job with two parts that point at the same recording: the part before the hold music, and the part after it. Each part takes source_in and duration, the result is one gapless file, and you send that file to STT so you do not pay to transcribe music. The job is $0.01 flat.
The fields are from the timeline audio page: parts[] takes 1 to 20 items, each { url, source_in?, duration? }, and the join is sample-domain with no silence at the seams.
Say a 12-minute call (720 s) has hold music from 3:10 to 6:40. STT takes at most 10 minutes per request, so removing the 210 seconds of music also brings the file under the limit: the parts 0 to 190 seconds and 400 to 720 seconds give 510 seconds. The numbers below are an invented example.
| Part | source_in (s) | duration (s) |
|---|---|---|
| Before hold | 0 | 190 |
| After hold | 400 | 320 |
| Joined length | 510 s (8.5 min) |
The request body
The recording must already be audio on media.sume.com in your workspace; import it first with POST /v1/media-imports, which needs an Idempotency-Key.
{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/call.wav", "source_in": 0, "duration": 190 },
{ "url": "https://media.sume.com/artifacts/artf_demo/call.wav", "source_in": 400, "duration": 320 }
],
"output": { "format": "wav" }
}Mapping word times back
The result carries segments[] with the offset of each part in the new file. Keep them: a word that STT places at 200 seconds in the joined file sits at 410 seconds in the original, because it falls in the second part and 210 seconds were removed. Part two starts at 190 in the new file and at 400 in the original, a difference of 210 seconds.
Format and errors
Keep the format on wav. The docs say mp3 adds priming padding at every edge, which matters at a seam. Concat parts must share a channel layout, and the same file used twice does.
Cost
Total cost is the $0.01 concat job plus STT at $0.01 per audio minute on the joined file, so the 8.5 minutes above come to about 9 cents; the 3.5 minutes of music you skipped would have cost about 4 cents more, and the original would not fit in one request if it were longer than 10 minutes.
Finding the cut points
You need the start and end of the hold stretch. If you have a transcript, the gap in words[] is the clue: a run of seconds with no words, or with garbled ones, between two phone-line phrases. Otherwise listen once and note the times, since the cuts are only as accurate as your times.
Give yourself a small margin. Cut a few tenths of a second inside the music at both ends rather than inside speech, because a clipped first word is worse than half a second of music left in.
If a call has more than one hold, add more parts to the same job; a concat job takes up to 20. Each extra part is just another source_in and duration pair on the same URL, and the price stays $0.01 for the job.
- Take cut times from the word gaps, or by listening.
- Cut inside the music, not inside speech.
- Several holds mean several parts, up to 20.
Sources
Related posts
More in Developers
- Idempotency-Key from an order id and version, never a fresh uuid
A fresh uuid per request makes Idempotency-Key do nothing. Derive it from the order id plus a version you bump only to re-run. Scope: one Format, 255 chars.
- How do I narrate a DIY tutorial step by step with a TTS API?
Narrate an 8-step DIY tutorial with one TTS job per step: 1,570 characters, $0.10 on Sume. Why per-step jobs make a fixed step a 1-cent redo.
- Do I pay for a failed AI avatar video job? Refunds on Sume
Sume reserves the avatar video price at submit, captures it on completion, and releases or refunds it where a job fails. What it means for retries.
- Does PNG, JPEG or WebP change the price of an AI image on Sume?
No. On Sume's Image API, output_format picks the file type, not the price: per-image cards and GPT Image 2.5 token math ignore it. Which models list which.
Written by Sume