What is text-based video editing? How transcript cuts work

Text-based video editing cuts a video by deleting words from its transcript. How word timings become a cut list, and how to build it with Sume's API.

5 min readSume
All posts

Text-based video editing means editing a video by editing its transcript: you delete words or sentences from the text, and the matching stretches of picture and sound are cut from the video. It works because each transcribed word carries a start and an end time, so the words you keep become the time ranges you render.

With Sume's API you build that pipeline yourself: transcribe the clip with video inspect, turn the kept words into ranges, and render the ranges with Timeline 1.0. The facts come from the Video inspect, Timeline 1.0, and Audio detach docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code.

How does a transcript edit become a cut?

Every word in the transcript has a start and an end on the video's clock. When you delete words, the ones left form runs, and each run of kept words becomes one range, from its first word's start to its last word's end. Played back to back, those ranges are the edited video, and each boundary between them is a cut. The table maps each piece to the Sume field that carries it.

From Video inspect, Timeline 1.0, and the Sume API reference, read 2026-09-28.
PieceSume fieldWhat it holds
Wordstranscript.words[]word, start, and end, in seconds from the start of the video
Sentencestranscript.segments[]Gapless sentences with start and end, when you ask for segmentation.mode: "sentence"
Where a kept range begins in the clipvideo[].source_inThe in-point into the source file
Where it lands in the editvideo[].start, video[].durationIts start on the output and how long it plays
Its soundaudio.parts[]The same range cut from the clip's detached sound, up to 20 slices

How do I get a transcript with word timings?

Send the clip to POST /v1/video-inspect with transcribe: true; frames: false skips the stills. The clip must be your workspace's media.sume.com artifact or asset, and language_code is an optional hint (omit it for auto-detect). Add segmentation to edit by sentence. The transcript is billed at $0.01 per audio minute, reserved from the duration_seconds hint (one minute when you omit it), and the probe itself is unbilled:

curl -X POST https://api.sume.com/v1/video-inspect \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: talk-transcript-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "frames": false,
    "transcribe": true,
    "segmentation": { "mode": "sentence" }
  }'

How do I turn the kept words into ranges?

Your editing tool records which words were deleted; this function turns the rest into (source_in, duration) pairs. The API reference marks only word as required, so it skips entries without timings:

def kept_ranges(words, deleted):
    """(source_in, duration) for each run of kept words.
    words: transcript.words[]; deleted: indexes the editor removed."""
    ranges, run = [], None
    for i, w in enumerate(words):
        if i in deleted:
            run = None  # a deleted word ends the run
        elif "start" in w and "end" in w:
            if run is None:
                run = [w["start"], w["end"]]
                ranges.append(run)
            else:
                run[1] = w["end"]  # same run, later end
    return [(s, round(e - s, 3)) for s, e in ranges]

Where should each cut go?

  • Inside a pause. Leave a small margin past each kept word's end, stopping short of the next deleted word, so word edges aren't clipped; how to remove silence from a video shows the margin in code.
  • At a sentence edge, when you edit by sentence. The API reference calls segments[] gapless, so deleting a sentence removes exactly its span and the kept sentences meet at their edges.
  • Never so close together that a range lasts under 0.2 seconds, the shortest Timeline slot. Merge a range that short with a neighbor, or drop it.

How do I render the edited video?

Play the ranges back to back in one Timeline 1.0 render. Each range becomes a video[] slot of the original clip and a matching audio.parts[] slice of the clip's sound, detached with POST /v1/audio-detach, both with the range's source_in and duration. Remove part of a video by API builds that render step by step, including what to do past 20 ranges, and how to remove silence from a video makes the same kind of cut list from pauses instead of deleted words.

The render is listed at $0.10 per output minute on API pricing; each step bills on its own, plus a 5.5% agent fee by default.

What are the limits?

  • Sume's API supplies the transcript and the render. The screen where someone deletes words is yours to build.
  • The transcript's duration_seconds hint stops at 600 seconds (10 minutes), so cut a longer video into parts with video trim and transcribe each one.
  • A clip with no audio track fails with inspect_source_has_no_audio.
  • The cuts land only where the word timings put them. Play the edit through before you publish it.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume