Five sermon clips from one recording: transcript, trim, captions
Find five 45-second moments in a 25-minute sermon from its transcript, trim them, and burn captions. Calls and a $1.37 total at Sume's published rates.

Transcribe the sermon, choose five passages by their sentence times, trim them and caption each: $1.37 for a 25-minute recording at Sume's published rates. Speech to text is $0.01 per audio minute, so the transcript itself is $0.25.
Get a transcript you can search
Detach the audio in two ranges (the source cap and a 900-second output cap apply), then run speech to text. Detach has a mono 16 kHz shape meant for it. Ask for sentence segmentation so each line carries its own start and end; those are your cut points.
{
"video_url": "https://media.sume.com/artifacts/artf_demo/sermon.mp4",
"range": { "start": 900 },
"channels": "mono",
"sample_rate": 16000
}Trim and caption
Each pick becomes a video-trim with start and duration, followed by a caption job on the result. Captions are one job per video up to 60 seconds, so keep each cut at or under a minute.
curl -X POST https://api.sume.com/v1/video-trim \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: sermon-pick-3" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/sermon.mp4",
"start": 742.5,
"duration": 45
}'Cost by step
Speech to text reserves from a duration hint (one minute if you send none, and the hint tops out at 600 seconds). The $0.25 figure is the published per-minute rate times 25 minutes; confirm the captured amount on the job.
| Step | Quantity | Rate | Subtotal |
|---|---|---|---|
| Audio detach | 2 ranges | $0.01 per job | $0.02 |
| Speech to text | 25 min | $0.01 per audio minute | $0.25 |
| Video trim | 5 | $0.02 per job | $0.10 |
| Video captions | 5 | $0.20 per job | $1.00 |
| Total | $1.37 |
Review before posting
Auto transcripts mishear names and scripture references. Pass script_text to the caption job when you have the exact words, which aligns your text to the speech timing, and read each clip once before it goes out.
Choosing passages
Look for passages with a single idea and a clear first sentence, since a clip cannot lean on the minutes before it. The sentence segments from speech to text give each line a start and end, so you can cut exactly on a sentence boundary instead of listening for it. Keep clips to about a minute so that one caption job covers each, and trim a second or two of lead-in so the viewer starts mid-thought rather than on a pause. Share only what the speaker has agreed to publish.
Related posts
More in Use cases
- Fix wrong text on a product label with ideogram-v4.5 and references
ideogram-v4.5 edits the first image and takes up to four more as references, five in total. Send a label photo plus a type sample; cost per edit by quality.
- 40 Etsy listings, 40 product clips: seven video models priced at 5 s
Price a 40-listing Etsy video batch across seven Sume video models at five seconds each, with the first-frame request that animates a listing photo.
- Google Ads Shorts previews: render 9:16, 16:9 and 1:1 first
Shorts previews need a Video view or reach campaign with Shorts selected. Render all three Google sizes from one master with Sume Timeline for $0.30.
- Google Ads video rejects MP3, WAV and PCM: pair audio with a still
Google Ads video upload does not accept audio files such as MP3, WAV or PCM. Turn a radio spot into an MP4 with a still and your audio using Sume Timeline.
Written by Sume