Podcast quote clips end abruptly: STT boundary_lead_ms tail padding
Sume STT segmentation boundary_lead_ms (0 to 500 ms, default 70) sets how long a sentence's tail runs before the cut. Tune it, then cut with timeline audio.

You pull a 12-second quote from a podcast, and the clip stops the instant the last word's timestamp ends. The final consonant is clipped, or the room tone drops out like a cut tape. The word time was right. A cut exactly at a word's end sounds abrupt. Sume STT sentence segmentation has a field for the tail: segmentation.boundary_lead_ms.
What the field does
With segmentation: {"mode": "sentence"} an STT job also returns segments[] with index, text, start, end and duration_seconds. The segments are gapless, so each ends where the next begins. According to the Sume API reference, boundary_lead_ms is the number of milliseconds of lead carried past a sentence's last word before the next segment starts, and the next segment absorbs the pause. It takes 0 to 500 and defaults to 70, the same rule and default as TTS 1.0. The segments are time ranges over your audio_url. This step does not slice any audio.
Raise it, then listen
For clean studio speech, 70 ms is often fine. For a conversational show where sentences trail off, try 150 to 250 ms. A larger value moves each boundary later, so the clip keeps more of its own tail and the next clip starts later. Past 300 ms you begin to include the next speaker's breath. Run one 10-minute file at 70, 200 and 400 and listen to the same five ends. Each run costs about 10 cents at $0.01 per minute, so the test is 30 cents.
import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
def run(path, body, key=None):
h = {**H, **({"Idempotency-Key": key} if key else {})}
r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
r.raise_for_status()
job = r.json()["data"]["job"]["id"]
while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
time.sleep(3)
res = requests.get(f"{B}/jobs/{job}/result", headers=H)
res.raise_for_status()
return res.json()["data"]["result"]
res = run("/stt-1.0/transcribe", {"audio_url": os.environ["EP_URL"], "duration_seconds": 600,
"segmentation": {"mode": "sentence", "boundary_lead_ms": 200}})
quotes = [s for s in res["segments"] if 8 <= s["duration_seconds"] <= 15]
ranges = [{"start": s["start"], "end": s["end"]} for s in quotes[:20]]
print(len(quotes), "candidates")
print(ranges[:3])Cutting the audio
Send ranges to POST /v1/timeline-1.0/audio in split mode: 1 to 20 ranges per job, at a flat $0.01. The url must be workspace audio on media.sume.com, so import the episode first with POST /v1/media-imports. See the timeline audio docs. If you choose mp3 output, the encoder adds a little priming padding to each clip. Use the default wav when you will edit further.
Check
- Listen to the last second of every clip, not the middle.
- If one file still ends hard, raise the lead for that file only. Do not raise it globally for the whole archive.
- If a clip ends with the start of someone else's word, lower it.
words[]keeps its own unshifted times. The lead only moves the segment boundary.
Sources
Related posts
More in Media tools
- Product spec sheet to video: Wan 3.0 footage, exact specs via compose
Make a spec-sheet video without a model re-typing your numbers: generate the footage with Wan 3.0, then put your own spec card on screen with Timeline compose.
- Product turntable ad: which Sume video models take two frames
For a product turn, give a video model the front and back photo as first and last frames. Six Sume rows take both; Grok Imagine takes a first frame only.
- Put a promo code on screen in an AI product video with caption cues
Burn a Black Friday code into a silent clip with $0.20 caption cues, key the retry by SKU and code, then pull stills with video-frames to check the text.
- Proof frame for a 1080x1920 export: video frames PNG at 1920
Pull a lossless 1080x1920 still from your vertical export with video frames (format png, max_edge 1920) to check captions and crop before you upload.
Written by Sume