Webinar recording to shorts: an API pipeline with Sume
Turn a webinar recording into short clips with four calls: inspect and transcribe by sentence, pick moments, trim, caption. Limits and costs included.

To turn a webinar recording into shorts, run four steps: video-inspect with transcribe:true to get sentence segments, choose the moments you want, video-trim each one, then video-captions on each clip. Every step is an API job, so the whole pipeline can run unattended. YouTube's Made On YouTube 2026 announcements put conversational editing for Shorts front and centre; this is the same job for recordings that live outside YouTube's editor.
The four steps
The source must already be on media.sume.com. Sume's media tools fetch nothing from the open internet, and POST /v1/media-imports accepts only public TikTok or Instagram URLs (YouTube is refused with unsupported_platform), so a webinar file needs to be uploaded as an asset first.
| Step | Endpoint | Price and limit |
|---|---|---|
| Transcribe by sentence | POST /v1/video-inspect with transcribe:true, segmentation.mode:"sentence" | $0.01 per audio minute; source up to 1800 s |
| Cut each moment | POST /v1/video-trim | $0.02 per job; 0.2-900 s |
| Burn captions | POST /v1/video-captions | $0.20 per video up to 60 s |
| Docs | Video inspect, Video trim, Video captions |
Choosing the moments
The inspect result returns segments[] with start and end times per sentence. Selection is the creative step and the one an agent or a person does: read the transcript, pick three to ten windows of 20 to 50 seconds that stand alone, and write down each start and end. Each pick becomes a trim with start and end.
Check probe.has_audio before you ask for a transcript. A source with no audio returns inspect_source_has_no_audio. If you leave out duration_seconds, the reservation assumes one minute; the hint accepts up to 600 s, so pass the real length for a clip of that size or shorter.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: webinar-42-clip-1" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/webinar.mp4",
"start": 612.4,
"end": 651.0,
"output": {"width": 1080, "height": 1920, "fps": 30}
}'Keeping the batch tidy
Name every clip job with an Idempotency-Key built from the webinar id, the clip number and a version, so a retry never cuts the same moment twice. Run the trims in parallel, since each is independent, and caption only the clips you keep: captions are the largest line at $0.20 each, so review the trimmed clips first and drop weak ones before paying for styling. Keep the transcript segments alongside each clip; they are the wording source if you later pass words to the captions call.
Limits
A single inspect call covers a source up to 1800 s (30 minutes), so a 90-minute webinar must be split into parts of 30 minutes or less before you upload it, since trim has the same 1800 s source cap. The output conform on trim re-encodes to 1080x1920 but does not reframe a speaker: it changes dimensions, so check the framing of a 16:9 recording before you publish. Captions need the clip to be a public HTTPS URL for video_url. Estimate a run of ten clips at about $0.20 (trim) + $2.00 (captions) plus transcription, and confirm live rates in GET /v1/catalog.
Sources
Related posts
More in Use cases
- Wedding highlight video: keep the vows audio, cut the rest
Cut a 2-minute wedding highlight from guest clips: take vows audio with audio-detach, place clips over it in Timeline 1.0, and add a quiet music bed.
- YouTube AI disclosure: AI backdrop and own-voice clone examples
YouTube lists cloning your own voice and extending a backdrop as edits needing no AI disclosure. What else is on that list, and what is not.
- Does YouTube's AI label hurt reach or monetization?
YouTube's help page says disclosing AI content won't limit a video's audience or monetization eligibility. What it says about skipping disclosure instead.
- YouTube automatic captions not available: what to do instead
YouTube lists silence, overlapping speakers and poor audio as reasons auto captions fail. Burn your own with Sume cues or a script_text-aligned caption job.
Written by Sume