One 45-second voiceover into 15 and 30 second cuts with audio split

One Sume timeline-audio split job can slice a 45-second voiceover into a 15-second and a 30-second file for $0.01. Cut on a sentence end.

4 min readSume
All posts

Slice one recorded voiceover into two lengths with a single POST /v1/timeline-1.0/audio request using operation: "split" and two ranges. The job is priced at $0.01 flat, uses no re-synthesis, and returns each range as its own durable media.sume.com file.

The rules are in the Timeline audio docs (read 2026-10-06).

What does the request look like?

The required fields are operation: "split", a top-level url of Sume-hosted audio, and ranges[] with 1 to 20 entries, each { start, end } in seconds. You can leave end out and the range goes to the end of the file. Ranges can overlap, so a 15-second cut can sit inside the 30-second one. An Idempotency-Key is required.

{
  "operation": "split",
  "url": "https://media.sume.com/artifacts/artf_demo/voiceover.wav",
  "ranges": [{ "start": 0, "end": 15.2 }, { "start": 0, "end": 30.6 }]
}

Where should I cut?

Cut at the end of a sentence. With TTS timestamps.words on and segmentation.mode: "sentence", the result returns gapless sentence segments with a default boundary of 70 ms after the last word, so each sentence's end is a safe cut point. The cut numbers above are examples; use the ones from your own result.

What format comes back?

The output defaults to WAV (pcm_s16le, sample-exact). MP3 is smaller but adds priming padding again at every edge, so keep WAV when you will join or lip-sync the file.

Sources

More in Media tools

All Media tools posts

Written by Sume