TTS take starts late: trim the lead-in with words[0].start
Find where speech really begins with timestamps.words and cut the silent lead-in with one timeline audio split job. Python for the range, $0.01 per job.

Request timestamps.words with the take, read words[0].start, and split the file from just before that time with one timeline audio job. If the first word starts at 0.62 seconds, a split range starting at 0.57 drops the silent lead-in and keeps a 50 millisecond cushion.
Do the same at the tail with the last word's end. Whether a given take has silence at the edges is something to measure, not assume.
Why trim the edges?
Joined takes add up the quiet at each edge, and a dead half second at every sentence boundary sounds like hesitation. Sume's timeline audio join is sample-domain with no added silence at the seams, so the pauses you hear are the ones inside the takes. Trimming them gives you control of pacing.
How do I find the speech start?
With timestamps.words set to true, the completed TTS job carries words[] with monotonic start and end values in seconds. The first entry's start is the lead-in length. The last entry's end is where speech stops. Everything outside those two numbers is silence or breath.
How do I cut it?
Timeline audio split takes a top-level url and 1 to 20 ranges, each with a start and an optional end. Each range comes back as its own audio_url segment. Use the default wav output, since mp3 re-adds priming padding at every edge. The sketch computes the range with a pad of 50 milliseconds, and runs as written.
The job is $0.01 flat, the same as a concat.
| Pad | Effect |
|---|---|
| 0 s | Tight; can clip a soft consonant |
| 0.05 s | Starting point |
| 0.15 s | Natural breath at joins |
| 0.30 s | Obvious pause; use for scene changes |
PAD = 0.05 # seconds of room kept around the speech
words = [{"word": "Hello", "start": 0.62, "end": 1.0}, {"word": "there", "start": 1.05, "end": 1.4}]
first, last = words[0]["start"], words[-1]["end"]
rng = {"start": max(0, round(first - PAD, 3)), "end": round(last + PAD, 3)}
print("lead-in", first, "s; split range", rng)
body = {"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/line.wav",
"ranges": [rng]}
print(body["ranges"])
What should I do with the new file?
Use the segment's audio_url as the spine in a Timeline render or as a concat part. Segment times shift by the amount you cut, so if captions or video starts were timed against the original take, subtract the cut start from them. Listen to the first second once: if a breath is missing, widen the pad.
Sources
Related posts
More in Developers
- Test call audio for voice agents: Sume TTS at 8 kHz mu-law
Generate repeatable phone-quality test utterances for a voice agent with Sume TTS output_format: 8000 Hz, pcm_mulaw. Fields, limits and a runnable script.
- TTS voice.id: a UUID or a voi_ library id? What Sume accepts
Sume TTS voice.id takes a voice UUID or a voi_ library id. Any other shape fails with 400 invalid_voice_id before a job is queued or credits are reserved.
- TTS word timestamps to Timeline slide starts in Python
Call the Sume TTS Router with word timestamps, find the word that opens each slide, and build the Timeline video array of start times in a short Python script.
- Twilio Play 40-minute file limit vs Sume TTS 1,200-second cap
Twilio warns that a played file over 40 minutes can drop the call. Sume TTS fails audio past 1,200 seconds, so a hold message fits. Format and caching notes.
Written by Sume