Reuse one sentence twice: timeline audio split allows overlaps
Timeline audio split takes up to 20 ranges that may overlap, so one hook sentence can open and close a video without re-recording or re-synthesising it.

Sume timeline audio split accepts 1 to 20 ranges, and the docs state that ranges can overlap. Overlap is how you reuse a sentence: one range covers the first six seconds of the voice, another covers seconds 4 to 20, and each comes back as its own file with its own audio_url. You can then use the first range as a cold open and the second as the body, without a second take or any re-synthesis.
The request
operation is split, url is a top-level media.sume.com audio file, and ranges[] holds { start, end? } objects in seconds. Leave end out and the range runs to the end of the file. Do not send parts (audio_split_takes_no_parts). A range with end at or before start returns audio_range_end_before_start. output.format is wav by default; keep wav if you will join the files again, since mp3 adds priming padding at each edge.
{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/spine.wav",
"ranges": [
{ "start": 0, "end": 6 },
{ "start": 4, "end": 20 },
{ "start": 20 }
]
}What comes back
The result is kind: timeline_audio with segments[], and each segment has an audio_url, its index and its duration. Ranges 1 and 2 above share seconds 4 to 6. The whole job is $0.01 flat, regardless of the number of ranges, and the produced audio is capped at 1800 s.
| Item | Value |
|---|---|
| Ranges per call | 1 to 20 |
| Overlap | allowed |
| Open-ended range | omit end |
| Format | wav (default) or mp3 |
| Price | $0.01 flat per job |
| Output cap | 1800 s |
| Source for a talking-head video | audio detach first |
Using the pieces
Cold open plus body is the classic case. Place the six-second piece first on the Timeline spine through audio.parts[], then the longer piece after it, and re-base the video[] starts against the segment offsets. The parts list takes up to 20 gapless slices joined in the sample domain, so there is no silence at the seams. If the clip you started from is an MP4, run audio detach once to get the wav that feeds the split.
Keep one thing in mind. The repeated sentence repeats on screen as well as in sound, so pick a line that bears repeating, and match the picture to it. A hook that appears twice with different footage reads as deliberate; the same footage twice reads as a mistake.
Sources
Related posts
More in Media tools
- Do fades shift my cuts? Timeline start times with transitions
In Timeline 1.0 a slot's declared start is its place on the audio spine, fade or no fade. The compiler compensates for the xfade overlap; gaps hold a frame.
- Two-voice dialogue audio on Sume: TTS lines joined with audio concat
Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.
- One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.
- video-filter dim amount 0 or 1.2 is refused: the (0, 1] range
Dim amount takes values above 0 up to 1. Zero, negatives and 1.2 return video_filter_amount_out_of_range. What it does, and how to lift a dark clip.
Written by Sume