Argil subtitles Top/Middle/Bottom and sizes vs Sume caption design
Argil subtitles take a styleId, a position and a size. Sume burns captions with a named style plus design overrides. Compare the knobs before you pick.

Argil's create-video endpoint turns subtitles on with a boolean and then lets you pick a styleId, a position of Top, Middle or Bottom, and a size of Small, Medium or Large. Sume burns captions with a named style such as slam, plus an optional design object that overrides colors, typography, placement, phrasing and motion for one request.
Argil's side is from its Create a new Video page; Sume's is from Video captions and Generate avatar video. I read all three on 2026-10-02.
What can you set on Argil subtitles?
The subtitles object has enable (required), an optional styleId, a position and a size. Position and size are three-value enums. That is a small surface: you choose a saved style and nudge it, and the page I read does not list finer controls such as colors or words per line.
What can you set on Sume captions?
Inline avatar captions take enabled, style (default slam), an optional font, a language hint and script_text. The styles are slam, punch, tiktok-green and korean-ad, plus the Hangul identities weight-shift, black-outline, highlight, pill-karaoke, clip-wipe and editorial-emphasis.
For finer control, standalone video captions accept design, which merges over the style: colors (base, active, stroke, accent), typography (font_size_ratio, stroke_width_px and others), placement (anchor_ratio, landscape_anchor_ratio), phrasing (max_words, max_chars, pause_seconds) and motion. Numbers outside range are a 400 at request time instead of a wrong render. design is not supported on punch or tiktok-green.
Comparison
Placement is where the two differ most: Argil gives three positions, Sume gives a numeric line center as a fraction of frame height.
| Control | Argil | Sume |
|---|---|---|
| Turn on | subtitles.enable | captions.enabled, or POST /v1/video-captions |
| Look | styleId | style name, optional design tokens |
| Position | Top, Middle, Bottom | placement.anchor_ratio (number) |
| Size | Small, Medium, Large | typography.font_size_ratio (number) |
| Existing video | Not on the page read | POST /v1/video-captions with a public video_url |
What are the limits on the Sume side?
Inline captions are rejected when the estimated duration is above 60 seconds. A caption failure soft-fails: the avatar job can still succeed with a clean video_url and captions.status=failed, and inline captions do not create a separate billed caption job. A Korean script with a Latin style such as slam is rejected with 400 caption_hangul_text_latin_style, so pick a Hangul style for Korean speech.
Sume burns captions into the picture. It does not export an SRT: the docs say SRT uploads are unsupported and that you pass phrase-level cues or segments instead. If your pipeline needs a sidecar subtitle file, that is a gap. Argil's docs index lists VTT and ASS export, which I did not verify on a page for this post.
Sources
Related posts
More in Comparisons
- Argil video status IDLE to DONE vs Sume queued to completed
Argil videos move through IDLE, GENERATING_AUDIO, GENERATING_VIDEO, DONE or FAILED; Sume jobs go queued, processing, completed. How to map polling code.
- Argil VIDEO_GENERATION_SUCCESS webhook vs Sume job.completed payload
Argil sends four webhook events with videoUrl in data; Sume sends signed terminal job.completed, job.failed and job.canceled events. Handler differences.
- AssemblyAI Universal-3.5 Pro $0.21 per hour vs Sume STT
AssemblyAI lists Universal-3.5 Pro async at $0.21 an hour plus $0.02 for speaker labels. Sume STT is $0.60 an hour with no add-on line. What to compare.
- Asset Studio 1-Click A/B Testing vs a Sume bulk run of hook variants
Google's Asset Studio adds Gemini Omni video and 1-Click A/B Testing. If your Q4 test spans channels, queue the hook variants as a Sume bulk run instead.
Written by Sume