Threads podcast transcript post: burn the quote onto the clip

For a podcast quote clip, burn captions from your exact transcript text with Sume video captions script_text, or authored cues. SRT uploads are not accepted.

4 min readSume
All posts

Threads' podcast toolkit includes transcript posts for sharing highlights. If you also want the quote burned onto a video clip, Sume POST /v1/video-captions can do it from the exact wording you supply: script_text aligns the burned-in text to your script, and cues burn authored text with no speech-to-text.

Threads details are from Meta's announcement; caption behavior from Video captions, read 2026-10-01.

What does Threads offer for transcripts?

Meta's post says you can share standout moments and featured guests with transcript, video, and guest card posts. It does not describe how a transcript post is built, so this page covers only the video side: a clip whose on-screen words match your transcript.

How do I keep the wording exact?

Pass script_text. The docs say Sume keeps the speech-to-text word timings as the timing source of truth and aligns burned-in wording to your script. That needs audible speech in the clip, and alignment can fail with typed errors such as script_alignment_mismatch; see the mismatch fix.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: quote-clip-001" \
  -d '{
    "video_url": "https://example.com/quote-clip.mp4",
    "script_text": "Your exact quote, as written in the transcript."
  }'

Which input fits which case?

Caption inputs from the Sume docs, read 2026-10-01.
InputBehavior
NoneSpeech-to-text transcribes the clip
script_textAligns burned-in wording to your script
cues or segmentsAuthored overlay with text, start, end; skips speech-to-text
wordsAlso a fixed-copy input

Can I upload an SRT file?

No. The constraints say SRT uploads and provider task ids are unsupported; pass phrase-level text as cues or segments. script_text, words, cues and segments are mutually exclusive. A silent clip fails as caption_no_speech, so use cues for it, as in silent clips with overlay cues.

What must the video URL look like?

A fetchable public HTTPS video URL. Localhost, private-network, non-HTTPS and signed or private URLs are rejected, so a clip you cut earlier must be reachable that way. The input is an existing finished clip; captions are burned in, not delivered as a sidecar.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume