TTS take starts late: trim the lead-in with words[0].start

Find where speech really begins with timestamps.words and cut the silent lead-in with one timeline audio split job. Python for the range, $0.01 per job.

4 min readSume
All posts

Request timestamps.words with the take, read words[0].start, and split the file from just before that time with one timeline audio job. If the first word starts at 0.62 seconds, a split range starting at 0.57 drops the silent lead-in and keeps a 50 millisecond cushion.

Do the same at the tail with the last word's end. Whether a given take has silence at the edges is something to measure, not assume.

Why trim the edges?

Joined takes add up the quiet at each edge, and a dead half second at every sentence boundary sounds like hesitation. Sume's timeline audio join is sample-domain with no added silence at the seams, so the pauses you hear are the ones inside the takes. Trimming them gives you control of pacing.

How do I find the speech start?

With timestamps.words set to true, the completed TTS job carries words[] with monotonic start and end values in seconds. The first entry's start is the lead-in length. The last entry's end is where speech stops. Everything outside those two numbers is silence or breath.

How do I cut it?

Timeline audio split takes a top-level url and 1 to 20 ranges, each with a start and an optional end. Each range comes back as its own audio_url segment. Use the default wav output, since mp3 re-adds priming padding at every edge. The sketch computes the range with a pad of 50 milliseconds, and runs as written.

The job is $0.01 flat, the same as a concat.

Pad choices for the trim (our suggestions, read 2026-10-02).
PadEffect
0 sTight; can clip a soft consonant
0.05 sStarting point
0.15 sNatural breath at joins
0.30 sObvious pause; use for scene changes
PAD = 0.05  # seconds of room kept around the speech

words = [{"word": "Hello", "start": 0.62, "end": 1.0}, {"word": "there", "start": 1.05, "end": 1.4}]
first, last = words[0]["start"], words[-1]["end"]
rng = {"start": max(0, round(first - PAD, 3)), "end": round(last + PAD, 3)}
print("lead-in", first, "s; split range", rng)
body = {"operation": "split",
        "url": "https://media.sume.com/artifacts/artf_demo/line.wav",
        "ranges": [rng]}
print(body["ranges"])

What should I do with the new file?

Use the segment's audio_url as the spine in a Timeline render or as a concat part. Segment times shift by the amount you cut, so if captions or video starts were timed against the original take, subtract the cut start from them. Listen to the first second once: if a breath is missing, widen the pad.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume