Split a voiceover into per-scene WAV files with Timeline ranges

One narration take, one Timeline audio split call: send up to 20 ranges and get a wav per scene plus offsets for $0.01, instead of a TTS take per scene.

5 min readSume
All posts

To cut one narration file into scene-length pieces, call POST /v1/timeline-1.0/audio with operation: "split", the file url, and up to 20 ranges[]. Each range becomes its own audio file in segments[]. The default output is sample-exact WAV, and a job costs $0.01 flat. The final range can leave end off to run to the end.

The request

Take the narration URL from your TTS job and the scene boundaries from its word or sentence timestamps.

curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: timeline-audio-split-001" \
  -d '{
    "operation": "split",
    "url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
    "ranges": [{ "start": 0, "end": 12.4 }, { "start": 12.4 }]
  }'

What comes back

The result is kind: timeline_audio with segments[], and each segment has its own audio_url, plus offsets you can use to place visuals. The output cannot exceed 1,800 seconds.

Replace the demo URL with your own file. Keep the output as wav if the pieces will be re-joined or if a piece will drive a lip-sync render, since mp3 output adds priming padding at each edge.

Where the boundaries come from

Ask the TTS job for timestamps: {words: true} and sentence segmentation. The sentence segments give you start and end times for each line. Group sentences into scenes and use the group's first start and last end as a range.

If you want one file per sentence straight from TTS, sentence segmentation with boundary_lead_ms already returns slices for wav or raw. Use the split route when your scenes are made of several sentences, or when the audio did not come from TTS.

Limits and costs

More than 20 scenes means more than one split job. Cut at a scene boundary so no piece straddles two jobs.

Timeline audio split limits from apps/docs models/timeline-audio.md (read 2026-10-05)
ItemValue
Ranges per job1 to 20
Price$0.01 flat per job
Output lengthUp to 1,800 s
Default formatwav, pcm_s16le

When this beats regenerating

Splitting one take keeps the voice, pacing and breath the same across scenes, because it is one performance. Regenerating each scene gives a fresh performance per scene, which can drift. For a 20-scene ad, one TTS take plus one $0.01 split is also cheaper in calls and easier to review.

If the narration came from a video, extract the audio with audio detach first, then split.

A small example

For a 30-second ad with three scenes, ask TTS for word timestamps, find the end of the last word in each scene, and send ranges such as 0 to 9.8, 9.8 to 21.1, and 21.1 onward. The three returned files line up exactly because the source was a wav.

Place each scene's visual with the matching offset from segments[].

Related posts

More in Media tools

All Media tools posts

Written by Sume