Trim silence from a TTS file: word timestamps and Timeline audio

Cut the lead-in and tail of a Sume TTS file by reading words[] start and end, then slice it with Timeline audio split for $0.01 or use source_in in a render.

5 min readSume
All posts

To trim silence from a Sume text-to-speech file, request timestamps.words, read the first word's start and the last word's end from the finished job, and cut the audio to that range. Two Sume tools do the cutting: Timeline audio split with a {start, end} range makes a new file for $0.01 flat, and Timeline's audio.source_in skips the lead-in inside a render. Add a small pad, 50 ms below, so the first consonant and the last release are not clipped.

The word timings come from the Sume API reference, the split and source_in rules from Timeline audio and Timeline 1.0, and the polling steps from Jobs and results, read on 2026-10-03. I did not measure how much silence a given voice leaves, so the sample timings below are made up for the demonstration.

Where do the word times come from?

With timestamps.words: true, the completed job result carries words[] with start and end in seconds from the start of the audio, in order, plus duration_seconds. Everything before words[0].start is lead-in, and everything after words[-1].end is tail. Sentence segmentation already cuts each sentence at 70 ms after its last word by default, and boundary_lead_ms changes that from 0 to 500; it only emits per-sentence audio for wav or raw.

Which tool trims it?

Choose by whether you want a file or only a cut inside one video.

Trimming a TTS file with Sume tools, from the Timeline and Timeline audio docs, read 2026-10-03.
ToolFieldOutputPrice
Timeline audio splitranges[{start, end}], up to 20A durable file per range with its own audio_url$0.01 flat per job
Timeline audio concatparts[{url, source_in, duration}]One joined file$0.01 flat per job
Timeline renderaudio.source_in with one audio.urlAn MP4; length stays duration_seconds$0.10 per output minute, rounded up

How do I build the split request?

This script takes the words from a result, pads both ends, and prints the split body. Send it to POST /v1/timeline-1.0/audio with an Idempotency-Key, using a URL already on media.sume.com, then poll the job. It runs offline.

words = [
    {"word": "Welcome", "start": 0.42, "end": 0.81},
    {"word": "to", "start": 0.83, "end": 0.92},
    {"word": "the", "start": 0.94, "end": 1.02},
    {"word": "show", "start": 1.05, "end": 1.61},
]
AUDIO_URL = "https://media.sume.com/example/voiceover.wav"
PAD = 0.05

start = max(0.0, words[0]["start"] - PAD)
end = words[-1]["end"] + PAD
body = {
    "operation": "split",
    "url": AUDIO_URL,
    "ranges": [{"start": round(start, 3), "end": round(end, 3)}],
}
print(body)
print("kept", round(end - start, 3), "s; dropped", round(start, 3), "s of lead-in")

What should I watch for?

Format matters when you cut.

  • Request wav for the TTS job. Timeline audio writes wav by default, which is sample-exact, while mp3 re-adds priming padding at every edge and undoes some of your trim.
  • audio.source_in is for a single audio.url spine; with audio.parts[] it fails as audio_source_in_requires_single_spine. Set duration_seconds to the trimmed length, since that sets the output length.
  • A split range needs end greater than start, or the job fails with audio_range_end_before_start.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume