Captions app 10-minute captions vs Sume's 60-second inline cap
Captions processes videos up to 10 minutes. Sume's inline avatar captions reject estimates above 60 seconds; standalone video-captions is a separate job.

The Captions app release notes say caption processing supports videos up to 10 minutes. In Sume, the limit depends on the route: inline captions on an avatar video are rejected when the estimated duration is above 60 seconds, and the standalone caption job is a separate endpoint.
Captions facts are from its release notes (an undated list); Sume facts from Generate avatar video and Video captions, read 2026-10-01.
What is Sume's inline caption limit?
Inline captions are an optional captions object on the avatar talking-video request. They burn into the final MP4 after generation, and the docs state that an estimated duration above 60 seconds is rejected. Avatar scripts are already limited to a 4-60 second estimate, so the two limits meet at 60.
| Route | Documented ceiling |
|---|---|
| Captions app, caption processing | Up to 10 minutes (release notes) |
| Sume inline captions | Rejected above 60 s estimate |
Sume standalone POST /v1/video-captions | No maximum stated; priced at $0.20 for videos up to 60 s |
How do I caption an existing longer clip?
Use the standalone job. POST /v1/video-captions takes a public HTTPS video_url and returns a job-backed captioned video. The page prices a standalone job at $0.20 for videos up to 60 seconds and states no maximum length, so test with your clip and read the job's status before you rely on a length. For the long-clip pattern, see add captions to a long video.
Do inline captions cost extra?
Inline captions do not create a separate billed video-caption job, but the docs describe a fixed $0.20 caption add-on included in the avatar-video estimate when enabled. A caption-stage failure soft-fails, so the avatar job can still succeed with a clean video_url and captions.status=failed.
Sources
Related posts
More in Use cases
- Captions AI Twin from a video upload vs Sume's photo avatar
Captions can build an AI Twin from an uploaded video. Sume's avatar API takes a prompt, a profile or a public HTTPS reference image, but no video input.
- Captions Avatar Looks in Prompt to Video: the Sume handle workflow
Captions saves looks inside one avatar for Prompt to Video. For several looks in Sume, make one avatar_handle per look and render one job per final video.
- Captions horizontal 16:9 AI Edit vs Sume avatar aspect ratios
Captions AI Edit now outputs 16:9. Sume avatar videos accept 16:9 too, at 720p only, with a script or plan estimated at 4-60 seconds. 9:16 is the default.
- Captions app SRT export vs Sume burned-in caption cues
The Captions app can export an SRT file for external caption tracks. Sume burns captions into the MP4 from authored cues with text, start and end.
Written by Sume