Podcast quote clips end abruptly: STT boundary_lead_ms tail padding

Sume STT segmentation boundary_lead_ms (0 to 500 ms, default 70) sets how long a sentence's tail runs before the cut. Tune it, then cut with timeline audio.

4 min readSume
All posts

You pull a 12-second quote from a podcast, and the clip stops the instant the last word's timestamp ends. The final consonant is clipped, or the room tone drops out like a cut tape. The word time was right. A cut exactly at a word's end sounds abrupt. Sume STT sentence segmentation has a field for the tail: segmentation.boundary_lead_ms.

What the field does

With segmentation: {"mode": "sentence"} an STT job also returns segments[] with index, text, start, end and duration_seconds. The segments are gapless, so each ends where the next begins. According to the Sume API reference, boundary_lead_ms is the number of milliseconds of lead carried past a sentence's last word before the next segment starts, and the next segment absorbs the pause. It takes 0 to 500 and defaults to 70, the same rule and default as TTS 1.0. The segments are time ranges over your audio_url. This step does not slice any audio.

Raise it, then listen

For clean studio speech, 70 ms is often fine. For a conversational show where sentences trail off, try 150 to 250 ms. A larger value moves each boundary later, so the clip keeps more of its own tail and the next clip starts later. Past 300 ms you begin to include the next speaker's breath. Run one 10-minute file at 70, 200 and 400 and listen to the same five ends. Each run costs about 10 cents at $0.01 per minute, so the test is 30 cents.

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

res = run("/stt-1.0/transcribe", {"audio_url": os.environ["EP_URL"], "duration_seconds": 600,
          "segmentation": {"mode": "sentence", "boundary_lead_ms": 200}})
quotes = [s for s in res["segments"] if 8 <= s["duration_seconds"] <= 15]
ranges = [{"start": s["start"], "end": s["end"]} for s in quotes[:20]]
print(len(quotes), "candidates")
print(ranges[:3])

Cutting the audio

Send ranges to POST /v1/timeline-1.0/audio in split mode: 1 to 20 ranges per job, at a flat $0.01. The url must be workspace audio on media.sume.com, so import the episode first with POST /v1/media-imports. See the timeline audio docs. If you choose mp3 output, the encoder adds a little priming padding to each clip. Use the default wav when you will edit further.

Check

  • Listen to the last second of every clip, not the middle.
  • If one file still ends hard, raise the lead for that file only. Do not raise it globally for the whole archive.
  • If a clip ends with the start of someone else's word, lower it.
  • words[] keeps its own unshifted times. The lead only moves the segment boundary.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume