Descript-style delete-a-word editing with Sume STT and video trim
Descript edits video by editing its transcript. Sume has no such editor, but STT word timings, video trim at $0.02 a cut and Timeline can rebuild the result.

Sume does not have a transcript editor, but you can build the core of one: transcribe, decide which words to drop, turn the kept words into time ranges, cut each range with video trim, and join the cuts with Timeline. Descript is the better tool if you want to do this by hand in an app; the recipe is for a pipeline that must run without a person.
Descript's own Help Center, read 2026-10-10, describes transcript-based editing along with Studio Sound, captions, text to speech, translation and lip-sync, and its Underlord assistant. Sume's side comes from the Timeline audio page, the video trim docs and Timeline 1.0.
The pieces on the Sume side
Speech to text is POST /v1/stt-1.0/transcribe with a public HTTPS audio_url. Word timings are always returned as words[], each {word, start, end} in seconds from the audio start. The optional duration_seconds accepts 1 to 600 and is described as a hint for usage reservation; leave it out and one minute is reserved.
Cutting is POST /v1/video-trim with video_url, a start, and one of end or duration. It takes a [start, end) range of one workspace clip and returns a new MP4, for $0.02 per job per the docs, with output between 0.2 and 900 seconds. precision: exact is frame-accurate; keyframe copies the stream and can start early. Joining uses Timeline 1.0, where each trimmed clip becomes a video[] slot with source_in set to 0.
| Step | In Descript | In a Sume pipeline |
|---|---|---|
| Get a transcript | Automatic in the app | POST /v1/stt-1.0/transcribe, words[] with times |
| Delete a word | Edit the text | Your code drops word indexes |
| Cut the video | The app updates the edit | One POST /v1/video-trim per kept range, $0.02 each |
| Join the pieces | Part of the project | POST /v1/timeline-1.0/render with one slot per cut |
| Review | Play it back in the app | You build it, or read the output file |
Turn kept words into ranges
The one part you have to write is the arithmetic. Consecutive kept words that are close together should become one range, so you pay for fewer cuts. This function does that, with a small pad so cuts do not clip a consonant.
def keep_ranges(words, deleted, pad=0.05, gap=0.35):
"""words: [{'word','start','end'}]; deleted: set of word indexes."""
ranges = []
for i, w in enumerate(words):
if i in deleted:
continue
start = max(0.0, w["start"] - pad)
end = w["end"] + pad
if ranges and start - ranges[-1][1] <= gap:
ranges[-1][1] = end
else:
ranges.append([start, end])
return [(round(a, 3), round(b, 3)) for a, b in ranges]
words = [
{"word": "so", "start": 0.0, "end": 0.3},
{"word": "um", "start": 0.4, "end": 0.7},
{"word": "hello", "start": 2.0, "end": 2.5},
{"word": "there", "start": 2.55, "end": 3.0},
]
print(keep_ranges(words, {1}))What it costs and where it breaks
Each kept range is one trim job. Ten ranges is ten trims, or $0.20 at the documented $0.02 rate, plus the Timeline render at $0.10 per output minute. Check live rates in GET /v1/catalog. Merging ranges with the gap argument lowers the count, but it also keeps any filler that sits inside the gap, so pick the gap to match how tightly you want to cut.
Be clear about the limits. The trim source must be one workspace clip of at most 1800 seconds, and outside URLs are refused, so import first. Trimmed pieces start at source_in 0 in the Timeline. Cuts made with exact precision re-encode, which takes more worker time than a stream copy. Nothing here fixes audio quality, removes background noise or edits the transcript visually; for those, use Descript's app.
Sources
Related posts
More in Use cases
- Dog groomer holiday booking clip from one photo: 6 s, $0.75
How a dog groomer can fill December slots with a 6-second vertical clip made from one grooming photo, plus a music bed. Sume prices to the cent.
- ElevenLabs dubbing editor in maintenance mode: redo one line on Sume
ElevenLabs says its dubbing editor gets critical fixes only. If you need per-line regeneration, a Sume TTS job per sentence gives you one retry per line.
- Escape room holiday team-event teaser: 10 s with sound for $1.25
An escape room can promote office holiday bookings with Gemini Omni Flash 1.1: a 10-second 720p clip with synced sound is $1.25. Limits and price table.
- Family holiday greeting: a 10-second Omni Flash card with music
Make a 10-second holiday greeting with Gemini Omni Flash 1.1 at 720p, a Lyria 3.5 music bed and a caption cue for the message. Chain and cost, read 2026-10-10.
Written by Sume