What is text-based video editing? How transcript cuts work
Text-based video editing cuts a video by deleting words from its transcript. How word timings become a cut list, and how to build it with Sume's API.

Text-based video editing means editing a video by editing its transcript: you delete words or sentences from the text, and the matching stretches of picture and sound are cut from the video. It works because each transcribed word carries a start and an end time, so the words you keep become the time ranges you render.
With Sume's API you build that pipeline yourself: transcribe the clip with video inspect, turn the kept words into ranges, and render the ranges with Timeline 1.0. The facts come from the Video inspect, Timeline 1.0, and Audio detach docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.
How does a transcript edit become a cut?
Every word in the transcript has a start and an end on the video's clock. When you delete words, the ones left form runs, and each run of kept words becomes one range, from its first word's start to its last word's end. Played back to back, those ranges are the edited video, and each boundary between them is a cut. The table maps each piece to the Sume field that carries it.
| Piece | Sume field | What it holds |
|---|---|---|
| Words | transcript.words[] | word, start, and end, in seconds from the start of the video |
| Sentences | transcript.segments[] | Gapless sentences with start and end, when you ask for segmentation.mode: "sentence" |
| Where a kept range begins in the clip | video[].source_in | The in-point into the source file |
| Where it lands in the edit | video[].start, video[].duration | Its start on the output and how long it plays |
| Its sound | audio.parts[] | The same range cut from the clip's detached sound, up to 20 slices |
How do I get a transcript with word timings?
Send the clip to POST /v1/video-inspect with transcribe: true; frames: false skips the stills. The clip must be your workspace's media.sume.com artifact or asset, and language_code is an optional hint (omit it for auto-detect). Add segmentation to edit by sentence. The transcript is billed at $0.01 per audio minute, reserved from the duration_seconds hint (one minute when you omit it), and the probe itself is unbilled:
curl -X POST https://api.sume.com/v1/video-inspect \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: talk-transcript-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"frames": false,
"transcribe": true,
"segmentation": { "mode": "sentence" }
}'How do I turn the kept words into ranges?
Your editing tool records which words were deleted; this function turns the rest into (source_in, duration) pairs. The API reference marks only word as required, so it skips entries without timings:
def kept_ranges(words, deleted):
"""(source_in, duration) for each run of kept words.
words: transcript.words[]; deleted: indexes the editor removed."""
ranges, run = [], None
for i, w in enumerate(words):
if i in deleted:
run = None # a deleted word ends the run
elif "start" in w and "end" in w:
if run is None:
run = [w["start"], w["end"]]
ranges.append(run)
else:
run[1] = w["end"] # same run, later end
return [(s, round(e - s, 3)) for s, e in ranges]Where should each cut go?
- Inside a pause. Leave a small margin past each kept word's
end, stopping short of the next deleted word, so word edges aren't clipped; how to remove silence from a video shows the margin in code. - At a sentence edge, when you edit by sentence. The API reference calls
segments[]gapless, so deleting a sentence removes exactly its span and the kept sentences meet at their edges. - Never so close together that a range lasts under 0.2 seconds, the shortest Timeline slot. Merge a range that short with a neighbor, or drop it.
How do I render the edited video?
Play the ranges back to back in one Timeline 1.0 render. Each range becomes a video[] slot of the original clip and a matching audio.parts[] slice of the clip's sound, detached with POST /v1/audio-detach, both with the range's source_in and duration. Remove part of a video by API builds that render step by step, including what to do past 20 ranges, and how to remove silence from a video makes the same kind of cut list from pauses instead of deleted words.
The render is listed at $0.10 per output minute on API pricing; each step bills on its own, plus a 5.5% agent fee by default.
What are the limits?
- Sume's API supplies the transcript and the render. The screen where someone deletes words is yours to build.
- The transcript's
duration_secondshint stops at 600 seconds (10 minutes), so cut a longer video into parts with video trim and transcribe each one. - A clip with no audio track fails with
inspect_source_has_no_audio. - The cuts land only where the word timings put them. Play the edit through before you publish it.
Sources
Related posts
More in Media tools
- Types of cuts in video editing: hard, jump, J, L, and more
The main types of cuts in video editing: hard, jump, J and L, match, cutaway, cross-cut, and smash cut. What each does, and how to build them.
- Vertical video resolution: 9:16 sizes from 720p to 4K
Vertical video resolution is 1080×1920 for 1080p, 720×1280 for 720p, and 2160×3840 for 4K. How to size any 9:16 frame, and what Sume can render.
- Video freezes but audio continues: causes and fixes
If a video freezes but the audio continues in every player, the file holds a still frame. How to check its streams, and the three Sume render causes.
- WAV vs MP3: which is better, and when to use each
WAV is exact and large; MP3 is small and lossy. Keep WAV while you edit or join audio, export MP3 last, and know that MP3 to WAV restores nothing.
Written by Sume