How do I turn one podcast episode into five quote clips for social?
Transcribe the episode in 10-minute chunks, split five quotes out, put the cover still under each and burn captions: $1.80 for a 28-minute episode on Sume.

To turn one podcast episode into five quote clips, transcribe it, pick five quotes, cut each with Timeline audio split, then render each one as a short video with your cover art as a still and burn captions on top. For a 28-minute episode, Sume charges $0.28 for the transcript, $0.01 to split it for transcription, $0.01 to split out the quotes, and $0.30 per clip for the render and the captions. That comes to $1.80 for five clips. Rates were read on 2026-10-07.
No generative video is needed: the picture is your own cover art, held still for the length of the quote. Plain audiograms are a normal format for podcasts, and they cost far less than animated ones.
Step 1: the transcript, in chunks
Import the episode audio with POST /v1/media-imports. One transcript job takes at most 10 minutes, so cut the episode into slices with operation: split on POST /v1/timeline-1.0/audio ($0.01, up to 20 ranges per job). A 28-minute episode is three slices of 10, 10 and 8 minutes, and three transcript jobs at $0.01 per audio minute: $0.28.
Send duration_seconds on each job (600, 600 and 480). Without it Sume reserves one minute. Add each slice's start offset (0, 600, 1200 seconds) to its word times, then merge the three into one transcript with episode-wide times.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ep-42-chunks" \
-d '{
"operation": "split",
"url": "https://media.sume.com/artifacts/artf_demo/episode42.wav",
"ranges": [
{ "start": 0, "end": 600 },
{ "start": 600, "end": 1200 },
{ "start": 1200 }
]
}'Step 2: pick quotes and cut them
Choose quotes of 20 to 45 seconds that make sense without the rest of the episode. Read the transcript, and use the word times to get each quote's start and end to the second. Then run one more split with five ranges, which returns segments[] with a durable audio_url for each. Ranges can overlap, which helps when two quotes share a sentence.
Leave a half second of room at each end, so the clip does not start on the middle of a word. Output stays wav by default, which is right for a file you will use again.
Step 3: render and caption
In Timeline 1.0, set each quote's audio as the render audio and place your cover art as a single still slot for the whole length. A still is a static hold, and the fit field (cover, contain, stretch or blur) controls how a square cover fills a vertical frame. A quote under a minute is one started minute, which is $0.10.
Then call POST /v1/video-captions with the rendered URL and the quote text as script_text, so the words are aligned to the speech. A caption job is $0.20 for a video up to 60 seconds. Check each clip on a phone: the cover should be clear, and the words should sit clear of the platform's own buttons.
Choosing quotes that work alone
A quote clip is seen by someone who has not heard the episode, so the best quotes are complete thoughts: a claim, a reason, and a stop. Skip anything that starts with 'and so' or refers to something said earlier. Prefer a quote where the speaker does not interrupt themselves, because a clean take is easier to caption.
Search the merged transcript for numbers, strong verbs and questions, since those make good opening lines. Then listen to your top ten and keep five. The transcript is a draft, so check each quote's words by ear before they go into script_text: a wrong name burned into a caption is worse than no caption.
Label each clip file with the episode number and the start time, such as ep42-0812, so you can find the source later. If a guest is in a clip, check that you have their permission to post it as a clip, not only the full episode.
Cost of the five clips
- The transcript is a one-time cost per episode. Five more quotes from the same episode cost only the split, the renders and the captions.
- If you do not need captions burned in, drop $1.00 and keep the audiogram.
- Keep the episode at 30 minutes or less per job: Timeline audio output is capped at 1,800 seconds, so cut a longer episode into halves first.
| Step | Calls | Unit price | Total |
|---|---|---|---|
| Split into transcript chunks | 1 | $0.01 | $0.01 |
| Transcribe 28 audio minutes | 3 | $0.01 per minute | $0.28 |
| Split five quotes | 1 | $0.01 | $0.01 |
| Render, one started minute each | 5 | $0.10 | $0.50 |
| Captions, up to 60 seconds each | 5 | $0.20 | $1.00 |
| Total | $1.80 |
Sources
Related posts
More in Use cases
- Post-call recap video with an AI avatar for prospects: script and cost
After a sales call, send a 30-second recap clip from a Sume avatar. Script structure, cost at standard, plus and max, and how to keep it honest and reviewed.
- Pre-rendered avatar greetings per visitor segment, not a live avatar
Instead of a live avatar for each visitor, render one short Sume avatar clip per segment ahead of time. Per-tier cost for six 12-second greetings.
- Product photo to a 6-second vertical ad: a still is a static hold
A still in a Timeline video slot is held, not animated: motion is accepted but ignored with a motion_ignored warning. Default 1080x1920, $0.10 for 6 seconds.
- Remove a date stamp from a scanned photo: crop first, then a mask edit
Remove an orange date stamp from a scanned family photo: crop it for free in Pillow, or mask it and edit with openai/gpt-image-2.5 from $0.0094 an image.
Written by Sume