Add captions to a recorded AI avatar call: what Sume needs

Tavus can record a live avatar call to your bucket. Sume's video captions need a public HTTPS URL and priced for clips up to 60 seconds. Steps and limits.

5 min readSume
All posts

To caption a recorded AI avatar call, get the recording to a public HTTPS URL and send it to Sume's POST /v1/video-captions with a style. Sume does not accept signed or private URLs for video inputs, and its documented caption price covers videos up to 60 seconds, so a long call needs trimming first. Sume does not join the call or caption it live.

The recording itself comes from the live avatar vendor. Tavus's create-conversation reference lists enable_recording, auto_start_recording and a recording_storage object for S3, GCS or Azure Blob (read 2026-10-03).

Why caption a recording at all?

Live avatar tools usually offer display subtitles during the call. Tavus lists enable_closed_captions as a property that enables subtitle display capability (read 2026-10-03). That is a viewer-side feature during the session; it does not give you an MP4 with the words burned in. Short vertical clips cut from a longer call, posted to feeds where people watch muted, need captions in the pixels.

What does Sume require of the video?

Three things from the docs. First, video_url must be a fetchable public HTTPS URL; localhost, private-network, non-HTTPS and signed or private URLs are rejected (Media inputs). Second, the clip needs audible speech, otherwise the job fails as caption_no_speech unless you pass your own cues. Third, the page prices each accepted caption job at $0.20 for videos up to 60 seconds under a fixed estimate, and says to confirm live pricing in GET /v1/catalog and the OpenAPI.

From a recorded avatar call to a captioned clip, read 2026-10-03
StepWhere it happensConstraint
Record the callThe live avatar vendor (Tavus enable_recording)Stored in your S3, GCS or Azure bucket
Pick the moment worth sharingYour editor or Sume video trimTrim takes a Sume-hosted clip, not an arbitrary URL
Make it reachableYour storagePublic HTTPS; signed or private URLs are rejected
Burn captionsPOST /v1/video-captionsAudible speech; documented price covers up to 60 s

What does the request look like?

Pass video_url, pick a style (slam is the default for Latin text), and optionally script_text if you have an accepted transcript; Sume keeps speech-to-text timing as the source of truth and aligns your wording to it. If alignment fails you get script_alignment_mismatch or script_alignment_failed, and the suggested fix is to simplify or omit script_text.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: call-clip-caption-001" \
  -d '{
    "video_url": "https://example.com/public/clip-from-call.mp4",
    "style": "slam",
    "language": "en"
  }'

Which caption style fits a call clip?

Without a style, wording decides: slam for Latin text and black-outline for Korean. You can also name punch or tiktok-green, and the Hangul identities for Korean speech. A Korean script on a Latin style is rejected with 400 caption_hangul_text_latin_style instead of burning unreadable boxes. design overrides colours, placement, phrasing and motion for one request (Video captions).

To try a second look on the same clip, pass source_caption_id instead of video_url. Sume reuses the source video and word timings, so no second speech-to-text runs, though the restyle is still billed as a render.

What can go wrong with recordings?

Three things. A presigned URL that is valid for you is rejected by Sume as a signed or private URL. A recording where the avatar is silent for the highlighted stretch fails as caption_no_speech unless you send authored cues. And a long call is priced and tested only up to 60 seconds, so submit the highlight, not the hour.

Also keep the legal side in view. Tavus's create-conversation reference includes a policy option for the EU AI Act, and a recording of a conversation features a real person's voice and face. Get consent for the recording and for the reuse before it becomes a public clip.

What about the long recording?

Sume's caption page documents its price for videos up to 60 seconds. For a longer recording, cut the highlight first. Sume's video trim takes one Sume-hosted clip plus a start and an end or duration, and the docs say to import media first through POST /v1/media-imports. Then caption the trimmed clip. Keep consent in mind: a recorded call has another person in it, so check that the people on it agreed to the recording and to reuse.

  • Public HTTPS URL only; presigned links will be rejected.
  • Audible speech, or authored cues for silent clips.
  • One job per clip of up to 60 seconds.
  • Restyle without re-transcribing by passing source_caption_id.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume