Descript-style delete-a-word editing with Sume STT and video trim

Descript edits video by editing its transcript. Sume has no such editor, but STT word timings, video trim at $0.02 a cut and Timeline can rebuild the result.

6 min readSume
All posts

Sume does not have a transcript editor, but you can build the core of one: transcribe, decide which words to drop, turn the kept words into time ranges, cut each range with video trim, and join the cuts with Timeline. Descript is the better tool if you want to do this by hand in an app; the recipe is for a pipeline that must run without a person.

Descript's own Help Center, read 2026-10-10, describes transcript-based editing along with Studio Sound, captions, text to speech, translation and lip-sync, and its Underlord assistant. Sume's side comes from the Timeline audio page, the video trim docs and Timeline 1.0.

The pieces on the Sume side

Speech to text is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url. Word timings are always returned as words[], each {word, start, end} in seconds from the audio start. The optional duration_seconds accepts 1 to 600 and is described as a hint for usage reservation; leave it out and one minute is reserved.

Cutting is POST /v1/video-trim with video_url, a start, and one of end or duration. It takes a [start, end) range of one workspace clip and returns a new MP4, for $0.02 per job per the docs, with output between 0.2 and 900 seconds. precision: exact is frame-accurate; keyframe copies the stream and can start early. Joining uses Timeline 1.0, where each trimmed clip becomes a video[] slot with source_in set to 0.

Descript capability from help.descript.com (read 2026-10-10); Sume routes and rates from the Sume docs.
StepIn DescriptIn a Sume pipeline
Get a transcriptAutomatic in the appPOST /v1/stt-1.0/transcribe, words[] with times
Delete a wordEdit the textYour code drops word indexes
Cut the videoThe app updates the editOne POST /v1/video-trim per kept range, $0.02 each
Join the piecesPart of the projectPOST /v1/timeline-1.0/render with one slot per cut
ReviewPlay it back in the appYou build it, or read the output file

Turn kept words into ranges

The one part you have to write is the arithmetic. Consecutive kept words that are close together should become one range, so you pay for fewer cuts. This function does that, with a small pad so cuts do not clip a consonant.

def keep_ranges(words, deleted, pad=0.05, gap=0.35):
    """words: [{'word','start','end'}]; deleted: set of word indexes."""
    ranges = []
    for i, w in enumerate(words):
        if i in deleted:
            continue
        start = max(0.0, w["start"] - pad)
        end = w["end"] + pad
        if ranges and start - ranges[-1][1] <= gap:
            ranges[-1][1] = end
        else:
            ranges.append([start, end])
    return [(round(a, 3), round(b, 3)) for a, b in ranges]

words = [
    {"word": "so", "start": 0.0, "end": 0.3},
    {"word": "um", "start": 0.4, "end": 0.7},
    {"word": "hello", "start": 2.0, "end": 2.5},
    {"word": "there", "start": 2.55, "end": 3.0},
]
print(keep_ranges(words, {1}))

What it costs and where it breaks

Each kept range is one trim job. Ten ranges is ten trims, or $0.20 at the documented $0.02 rate, plus the Timeline render at $0.10 per output minute. Check live rates in GET /v1/catalog. Merging ranges with the gap argument lowers the count, but it also keeps any filler that sits inside the gap, so pick the gap to match how tightly you want to cut.

Be clear about the limits. The trim source must be one workspace clip of at most 1800 seconds, and outside URLs are refused, so import first. Trimmed pieces start at source_in 0 in the Timeline. Cuts made with exact precision re-encode, which takes more worker time than a stream copy. Nothing here fixes audio quality, removes background noise or edits the transcript visually; for those, use Descript's app.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume