Five sermon clips from one recording: transcript, trim, captions

Find five 45-second moments in a 25-minute sermon from its transcript, trim them, and burn captions. Calls and a $1.37 total at Sume's published rates.

3 min readSume
All posts

Transcribe the sermon, choose five passages by their sentence times, trim them and caption each: $1.37 for a 25-minute recording at Sume's published rates. Speech to text is $0.01 per audio minute, so the transcript itself is $0.25.

Get a transcript you can search

Detach the audio in two ranges (the source cap and a 900-second output cap apply), then run speech to text. Detach has a mono 16 kHz shape meant for it. Ask for sentence segmentation so each line carries its own start and end; those are your cut points.

{
  "video_url": "https://media.sume.com/artifacts/artf_demo/sermon.mp4",
  "range": { "start": 900 },
  "channels": "mono",
  "sample_rate": 16000
}

Trim and caption

Each pick becomes a video-trim with start and duration, followed by a caption job on the result. Captions are one job per video up to 60 seconds, so keep each cut at or under a minute.

curl -X POST https://api.sume.com/v1/video-trim \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: sermon-pick-3" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/sermon.mp4",
    "start": 742.5,
    "duration": 45
  }'

Cost by step

Speech to text reserves from a duration hint (one minute if you send none, and the hint tops out at 600 seconds). The $0.25 figure is the published per-minute rate times 25 minutes; confirm the captured amount on the job.

Sume list price, read 2026-10-08
StepQuantityRateSubtotal
Audio detach2 ranges$0.01 per job$0.02
Speech to text25 min$0.01 per audio minute$0.25
Video trim5$0.02 per job$0.10
Video captions5$0.20 per job$1.00
Total$1.37

Review before posting

Auto transcripts mishear names and scripture references. Pass script_text to the caption job when you have the exact words, which aligns your text to the speech timing, and read each clip once before it goes out.

Choosing passages

Look for passages with a single idea and a clear first sentence, since a clip cannot lean on the minutes before it. The sentence segments from speech to text give each line a start and end, so you can cut exactly on a sentence boundary instead of listening for it. Keep clips to about a minute so that one caption job covers each, and trim a second or two of lead-in so the viewer starts mid-thought rather than on a pause. Share only what the speaker has agreed to publish.

Related posts

More in Use cases

All Use cases posts

Written by Sume