Webinar to Shorts: pick moments by transcript, not fixed cuts
Fixed-interval cuts make interchangeable Shorts. Read the webinar's sentence segments, choose complete thoughts, then trim and caption each with Sume.

To turn a webinar into Shorts that do not look mass-produced, choose the moments from the transcript instead of cutting every 60 seconds. Run Video inspect with transcribe: true and segmentation.mode: "sentence", read the sentence segments for complete thoughts, then trim each chosen range and caption it. A fixed slice starts and ends mid-sentence and gives every clip the same rhythm, which is the pattern YouTube's monetization policy calls content that feels interchangeable from video to video (read 2026-10-03).
Step 1: get sentence segments
Video inspect reads one clip on your workspace's media.sume.com. The webinar must be uploaded or imported first, and it must be 1800 seconds or shorter. With transcribe: true it runs Sume STT on the audio at $0.01 per audio minute; omit duration_seconds and the reservation is one minute, and the documented maximum hint is 600 seconds. frames: false skips the stills if you only want the transcript.
Step 2: choose the ranges
Each segment carries start and end times. Read them as candidate boundaries and choose spans that are one complete thought: a question and its answer, a claim and its example. Selection is an editorial or agent step, not an API call. If you use a Sume agent for it, give it the segments and a rule such as one point per clip and no clip over 60 seconds.
Step 3: trim and caption each range
Video trim takes video_url, start and exactly one of end or duration, and returns a new MP4. Frame-accurate precision: "exact" is the default. Then Video captions burns subtitles; leaving style omitted lets the wording decide the look, and you can pass your own cues for a per-clip hook line.
| Job | Rate | Notes |
|---|---|---|
| Transcribe once | $0.01 per audio minute | Per webinar, not per clip |
| Trim one clip | $0.02 per job | Output up to 900 seconds |
| Caption one clip | $0.20 per job | Video up to 60 seconds |
| Total per clip after transcript | $0.22 | Trim plus caption |
Where this stops
Sume does not choose the best moments for you in this flow; the transcript tells you what was said, not what was good. Confirm current rates in GET /v1/catalog. For the whole pipeline end to end, including delivery, see Webinar recording to Shorts.
Sources
Related posts
More in Media tools
- WebVTT cue settings (line, size) vs Sume anchor_ratio and width
WebVTT line:78%,center and size:90% have close matches in Sume's caption design fields. A converter script, and what per-cue settings a burned render loses.
- Where to break subtitle lines: BBC rules and Sume caption cues
The BBC says one sentence per subtitle and no article-noun splits. Sume sentence segments cover the first rule; cues with a line break cover the rest.
- Who is speaking in subtitles: BBC colours, dashes, and Sume cues
WCAG 1.2.2 wants speaker identification in captions. The BBC prefers colour, then dashes or labels. What Sume's burned-in captions can do and a dash script.
- YouTube Auto-sync captions vs Sume script_text: track or burned in
YouTube Auto-sync times your transcript into a caption track. Sume script_text aligns your wording into burned-in captions. What each needs and which to use.
Written by Sume