Add captions to a recorded AI avatar call: what Sume needs
Tavus can record a live avatar call to your bucket. Sume's video captions need a public HTTPS URL and priced for clips up to 60 seconds. Steps and limits.
To caption a recorded AI avatar call, get the recording to a public HTTPS URL and send it to Sume's POST /v1/video-captions with a style. Sume does not accept signed or private URLs for video inputs, and its documented caption price covers videos up to 60 seconds, so a long call needs trimming first. Sume does not join the call or caption it live.
The recording itself comes from the live avatar vendor. Tavus's create-conversation reference lists enable_recording, auto_start_recording and a recording_storage object for S3, GCS or Azure Blob (read 2026-10-03).
Why caption a recording at all?
Live avatar tools usually offer display subtitles during the call. Tavus lists enable_closed_captions as a property that enables subtitle display capability (read 2026-10-03). That is a viewer-side feature during the session; it does not give you an MP4 with the words burned in. Short vertical clips cut from a longer call, posted to feeds where people watch muted, need captions in the pixels.
What does Sume require of the video?
Three things from the docs. First, video_url must be a fetchable public HTTPS URL; localhost, private-network, non-HTTPS and signed or private URLs are rejected (Media inputs). Second, the clip needs audible speech, otherwise the job fails as caption_no_speech unless you pass your own cues. Third, the page prices each accepted caption job at $0.20 for videos up to 60 seconds under a fixed estimate, and says to confirm live pricing in GET /v1/catalog and the OpenAPI.
| Step | Where it happens | Constraint |
|---|---|---|
| Record the call | The live avatar vendor (Tavus enable_recording) | Stored in your S3, GCS or Azure bucket |
| Pick the moment worth sharing | Your editor or Sume video trim | Trim takes a Sume-hosted clip, not an arbitrary URL |
| Make it reachable | Your storage | Public HTTPS; signed or private URLs are rejected |
| Burn captions | POST /v1/video-captions | Audible speech; documented price covers up to 60 s |
What does the request look like?
Pass video_url, pick a style (slam is the default for Latin text), and optionally script_text if you have an accepted transcript; Sume keeps speech-to-text timing as the source of truth and aligns your wording to it. If alignment fails you get script_alignment_mismatch or script_alignment_failed, and the suggested fix is to simplify or omit script_text.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: call-clip-caption-001" \
-d '{
"video_url": "https://example.com/public/clip-from-call.mp4",
"style": "slam",
"language": "en"
}'Which caption style fits a call clip?
Without a style, wording decides: slam for Latin text and black-outline for Korean. You can also name punch or tiktok-green, and the Hangul identities for Korean speech. A Korean script on a Latin style is rejected with 400 caption_hangul_text_latin_style instead of burning unreadable boxes. design overrides colours, placement, phrasing and motion for one request (Video captions).
To try a second look on the same clip, pass source_caption_id instead of video_url. Sume reuses the source video and word timings, so no second speech-to-text runs, though the restyle is still billed as a render.
What can go wrong with recordings?
Three things. A presigned URL that is valid for you is rejected by Sume as a signed or private URL. A recording where the avatar is silent for the highlighted stretch fails as caption_no_speech unless you send authored cues. And a long call is priced and tested only up to 60 seconds, so submit the highlight, not the hour.
Also keep the legal side in view. Tavus's create-conversation reference includes a policy option for the EU AI Act, and a recording of a conversation features a real person's voice and face. Get consent for the recording and for the reuse before it becomes a public clip.
What about the long recording?
Sume's caption page documents its price for videos up to 60 seconds. For a longer recording, cut the highlight first. Sume's video trim takes one Sume-hosted clip plus a start and an end or duration, and the docs say to import media first through POST /v1/media-imports. Then caption the trimmed clip. Keep consent in mind: a recorded call has another person in it, so check that the people on it agreed to the recording and to reuse.
- Public HTTPS URL only; presigned links will be rejected.
- Audible speech, or authored
cuesfor silent clips. - One job per clip of up to 60 seconds.
- Restyle without re-transcribing by passing
source_caption_id.
Sources
Related posts
More in Media tools
- AI music generator for covers: what Sume cannot do, and a swap
Sume cannot cover an existing song: no audio input and no song-to-song path. Licensed cover platforms are still in development. Here is an original-track swap.
- Apple Podcasts transcript speaker names: VTT, and Sume STT gaps
Apple shows speaker names when you provide a VTT. Sume STT has no speaker labels, so transcribe each track and merge with voice tags. Script included.
- FLUX 21:9 ultrawide image: FLUX.2 on Sume vs FLUX 3 ratios
FLUX 3 Image lists ratios from 21:9 to 9:21. On Sume, FLUX.2 Pro and Flex take 21:9 and 9:21 too, from a 13-ratio list with no auto. A banner request in Python.
- Spotify podcast transcript upload: VTT, 5MB and Sume STT parts
Spotify takes VTT or SRT up to 5MB, with timestamps. Build one from Sume STT in 10-minute parts, stitch the cues, and upload from Spotify for Creators.
Written by Sume