Split a voiceover into per-scene WAV files with Timeline ranges
One narration take, one Timeline audio split call: send up to 20 ranges and get a wav per scene plus offsets for $0.01, instead of a TTS take per scene.

To cut one narration file into scene-length pieces, call POST /v1/timeline-1.0/audio with operation: "split", the file url, and up to 20 ranges[]. Each range becomes its own audio file in segments[]. The default output is sample-exact WAV, and a job costs $0.01 flat. The final range can leave end off to run to the end.
The request
Take the narration URL from your TTS job and the scene boundaries from its word or sentence timestamps.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: timeline-audio-split-001" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
"ranges": [{ "start": 0, "end": 12.4 }, { "start": 12.4 }]
}'What comes back
The result is kind: timeline_audio with segments[], and each segment has its own audio_url, plus offsets you can use to place visuals. The output cannot exceed 1,800 seconds.
Replace the demo URL with your own file. Keep the output as wav if the pieces will be re-joined or if a piece will drive a lip-sync render, since mp3 output adds priming padding at each edge.
Where the boundaries come from
Ask the TTS job for timestamps: {words: true} and sentence segmentation. The sentence segments give you start and end times for each line. Group sentences into scenes and use the group's first start and last end as a range.
If you want one file per sentence straight from TTS, sentence segmentation with boundary_lead_ms already returns slices for wav or raw. Use the split route when your scenes are made of several sentences, or when the audio did not come from TTS.
Limits and costs
More than 20 scenes means more than one split job. Cut at a scene boundary so no piece straddles two jobs.
| Item | Value |
|---|---|
| Ranges per job | 1 to 20 |
| Price | $0.01 flat per job |
| Output length | Up to 1,800 s |
| Default format | wav, pcm_s16le |
When this beats regenerating
Splitting one take keeps the voice, pacing and breath the same across scenes, because it is one performance. Regenerating each scene gives a fresh performance per scene, which can drift. For a 20-scene ad, one TTS take plus one $0.01 split is also cheaper in calls and easier to review.
If the narration came from a video, extract the audio with audio detach first, then split.
A small example
For a 30-second ad with three scenes, ask TTS for word timestamps, find the end of the last word in each scene, and send ranges such as 0 to 9.8, 9.8 to 21.1, and 21.1 onward. The three returned files line up exactly because the source was a wav.
Place each scene's visual with the matching offset from segments[].
Related posts
More in Media tools
- Stitching AI video clips in Sume Timeline: the rules that return 400
Timeline 1.0 joins model clips into one MP4: first slot at 0, import first, fades up to 1 s, at most 8 chained fades. Error codes, cost, and a curl request.
- Sume BGM catalog: which tracks need a credit line and which do not
Sume's own BGM loops need no third-party credit. The licensed km tracks are CC BY 4.0 and need the Kevin MacLeod credit on every public export.
- Timeline output size: 256 to 2160 per side, so no 3840x2160
Sume Timeline 1.0 takes even output width and height from 256 to 2160, default 1080x1920. What that means for 4K, 2560x1080 ultrawide and portrait masters.
- Sume Timeline transitions: 6 types, 1 s cap and the 8-fade chain limit
Timeline 1.0 has fade, wipeleft, wiperight, slideup, slidedown and dissolve. Duration is capped at 1 s and half the shorter clip; 9 chained fades fail.
Written by Sume