Redact names from a transcript and the audio: STT words, then split

Sume STT returns word timings with a type field. Find the names you list, blank them in text, and cut the same spans from the audio for $0.01.

5 min readSume
All posts

Sharing a recording usually means removing a few names, phone numbers or account ids first. A transcript with the names blanked out is easy. The audio is the part people forget: the voice still says the name. Sume gives you two pieces to solve both, STT word timings and an audio split endpoint. It does not decide what counts as sensitive. You supply the list.

Step 1: word timings

POST /v1/stt-1.0/transcribe always returns words[], each with word, start, end and type (word or spacing). The list is capped at 20,000 entries. If it is cut, words_truncated is true and words_total gives the real count, so a long file is never shortened silently. A job takes at most 10 minutes of audio, public HTTPS URL only.

Step 2: find the spans

Match your list against the tokens with type == "word", lower-cased and stripped of punctuation. Pad each span by about 0.15 seconds on each side so the sound of a clipped syllable does not leak. Merge spans that touch.

def spans(words, banned, pad=0.15):
    out = []
    for w in words:
        tok = w["word"].lower().strip(".,!?;:\"'")
        if w.get("type") == "word" and tok in banned:
            a, b = max(0, w["start"] - pad), w["end"] + pad
            if out and a <= out[-1][1]:
                out[-1][1] = b
            else:
                out.append([a, b])
    return out

def keep(spans, total):
    cuts, t = [], 0.0
    for a, b in spans:
        cuts.append({"start": t, "end": a}); t = b
    return [c for c in cuts if c["end"] > c["start"]] + [{"start": t, "end": total}]

Step 3: cut the audio

POST /v1/timeline-1.0/audio with parts (1 to 20 items, each url, source_in, duration) concatenates kept stretches into one file, capped at 1,800 seconds, at a flat $0.01 per job. The timeline audio docs show the shape. The URL must be workspace audio on media.sume.com, so import the file first with POST /v1/media-imports. Turn each kept range from keep into a part: source_in is start, duration is end - start. With more than 20 kept ranges, join in two passes.

Limits

  • This removes the word. It does not bleep or replace it. Silence will be audible, so tell listeners the file is edited.
  • Provider transcription misses names, which is when they are most likely to be spelled differently than your list. Add variants.
  • Check the result by transcribing the edited audio again: about one cent for a minute. If a banned word shows up, widen the pad.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume