Trim a voiceover take with Timeline audio: source_in and duration
Cut a Sume-hosted voiceover with POST /v1/timeline-1.0/audio: one part with source_in and duration, wav by default, $0.01 flat, Idempotency-Key required.

Send a concat request with a single part and set source_in (where to start, in seconds) and duration (how long to take). Sume returns a new gapless wav of just that slice for a flat $0.01 per job. The audio must already be hosted on Sume, such as a TTS result.
The request needs an Idempotency-Key header, so a retry returns the same job instead of a second charge.
The request
Set SUME_API_KEY and TAKE_URL (a Sume-hosted audio URL) in your shell, then run it.
curl -sS https://api.sume.com/v1/timeline-1.0/audio \
-H "x-api-key: $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: trim-take-4-v1" \
-d '{
"operation": "concat",
"parts": [
{ "url": "'"$TAKE_URL"'", "source_in": 1.2, "duration": 8.5 }
],
"output": { "format": "wav" }
}'Limits to know
The contract and docs set these bounds.
| Item | Value |
|---|---|
| Route | POST /v1/timeline-1.0/audio |
| Operations | concat (join parts) or split (cut ranges) |
| Parts or ranges per call | 1 to 20 |
| Per-part controls | Optional source_in (seconds, default 0) and duration |
| Source length | Up to 1,800 seconds |
| Output format | wav (default, sample-exact) or mp3 |
| Price | $0.01 flat per job |
| Mixed inputs | Parts must share one channel layout, else audio_parts_channel_mismatch |
| Hosting | All URLs must be Sume-hosted |
Why wav is the default
The reference explains that wav stays sample-exact, so the result can be joined again without picking up encoder delay, while mp3 adds priming padding at every edge. If you plan to trim, then join, then trim again, keep every intermediate in wav and convert to mp3 only on the last pass.
Cut and join in one call
Each part carries its own source_in and duration, so one call can pull the best sentence from three takes and join them in order. For ranges out of one long file, use operation: split with a url and up to 20 ranges, which return one file per range. That is a good fit for pulling clips out of a long TTS read.
The route works only on audio you already have on Sume. For the whole-video case, a Timeline 1.0 render can take the audio parts directly, so you can skip this job when the audio is only needed in one render.
Sources
Related posts
More in Media tools
- Trim, caption, compose: Sume returns new files, keeping your original
Sume's trim and caption tools return new MP4s and leave the source unchanged. Why that keeps the original AI generation intact, and what each step costs.
- Two-voice dialogue audio on Sume: TTS lines joined with audio concat
Build a role-play or interview track by generating one TTS line per turn in each speaker's voice, then joining up to 20 turns into one gapless file for $0.01.
- One lip-sync model per video: Fabric runs at 25 fps, measure H3 Max
Mixing Fabric and MiniMax H3 Max lip-sync clips in one Sume video risks a frame-rate mismatch. Pick one per run and check fps with ffprobe on the first clip.
- video-filter dim amount 0 or 1.2 is refused: the (0, 1] range
Dim amount takes values above 0 up to 1. Zero, negatives and 1.2 return video_filter_amount_out_of_range. What it does, and how to lift a dark clip.
Written by Sume