Trim silence from a TTS file: word timestamps and Timeline audio
Cut the lead-in and tail of a Sume TTS file by reading words[] start and end, then slice it with Timeline audio split for $0.01 or use source_in in a render.

To trim silence from a Sume text-to-speech file, request timestamps.words, read the first word's start and the last word's end from the finished job, and cut the audio to that range. Two Sume tools do the cutting: Timeline audio split with a {start, end} range makes a new file for $0.01 flat, and Timeline's audio.source_in skips the lead-in inside a render. Add a small pad, 50 ms below, so the first consonant and the last release are not clipped.
The word timings come from the Sume API reference, the split and source_in rules from Timeline audio and Timeline 1.0, and the polling steps from Jobs and results, read on 2026-10-03. I did not measure how much silence a given voice leaves, so the sample timings below are made up for the demonstration.
Where do the word times come from?
With timestamps.words: true, the completed job result carries words[] with start and end in seconds from the start of the audio, in order, plus duration_seconds. Everything before words[0].start is lead-in, and everything after words[-1].end is tail. Sentence segmentation already cuts each sentence at 70 ms after its last word by default, and boundary_lead_ms changes that from 0 to 500; it only emits per-sentence audio for wav or raw.
Which tool trims it?
Choose by whether you want a file or only a cut inside one video.
| Tool | Field | Output | Price |
|---|---|---|---|
| Timeline audio split | ranges[{start, end}], up to 20 | A durable file per range with its own audio_url | $0.01 flat per job |
| Timeline audio concat | parts[{url, source_in, duration}] | One joined file | $0.01 flat per job |
| Timeline render | audio.source_in with one audio.url | An MP4; length stays duration_seconds | $0.10 per output minute, rounded up |
How do I build the split request?
This script takes the words from a result, pads both ends, and prints the split body. Send it to POST /v1/timeline-1.0/audio with an Idempotency-Key, using a URL already on media.sume.com, then poll the job. It runs offline.
words = [
{"word": "Welcome", "start": 0.42, "end": 0.81},
{"word": "to", "start": 0.83, "end": 0.92},
{"word": "the", "start": 0.94, "end": 1.02},
{"word": "show", "start": 1.05, "end": 1.61},
]
AUDIO_URL = "https://media.sume.com/example/voiceover.wav"
PAD = 0.05
start = max(0.0, words[0]["start"] - PAD)
end = words[-1]["end"] + PAD
body = {
"operation": "split",
"url": AUDIO_URL,
"ranges": [{"start": round(start, 3), "end": round(end, 3)}],
}
print(body)
print("kept", round(end - start, 3), "s; dropped", round(start, 3), "s of lead-in")What should I watch for?
Format matters when you cut.
- Request wav for the TTS job. Timeline audio writes wav by default, which is sample-exact, while mp3 re-adds priming padding at every edge and undoes some of your trim.
audio.source_inis for a singleaudio.urlspine; withaudio.parts[]it fails asaudio_source_in_requires_single_spine. Setduration_secondsto the trimmed length, since that sets the output length.- A split range needs
endgreater thanstart, or the job fails withaudio_range_end_before_start.
Sources
Related posts
More in Developers
- Try-on double click: one Idempotency-Key per shopper and garment
Stop a double-clicked Try on button from running a Sume Format twice. Derive the Idempotency-Key from shopper, garment and version, and read the 200 replay.
- Typed try-on results: output_schema with a video URL and your SKU
Bind an output_schema to a Sume try-on run to get the video as a typed field, and keep your SKU outside it because identifiers do not round-trip.
- TTS from an accepted script: transcript_source instead of pasted text
Sume TTS can read an accepted script by script_revision_id and sentence_ids, not pasted text. MCP tools tts_source_get and tts_source_verify_spine support it.
- TTS volume 0.5 to 2: set narration gain before mixing with music
Sume TTS 1.0 generation_config.volume runs from 0.5 to 2 alongside speed 0.6 to 1.5. How to set narration level before you mix with a music bed.
Written by Sume