Argil subtitles Top/Middle/Bottom and sizes vs Sume caption design

Argil subtitles take a styleId, a position and a size. Sume burns captions with a named style plus design overrides. Compare the knobs before you pick.

5 min readSume
All posts

Argil's create-video endpoint turns subtitles on with a boolean and then lets you pick a styleId, a position of Top, Middle or Bottom, and a size of Small, Medium or Large. Sume burns captions with a named style such as slam, plus an optional design object that overrides colors, typography, placement, phrasing and motion for one request.

Argil's side is from its Create a new Video page; Sume's is from Video captions and Generate avatar video. I read all three on 2026-10-02.

What can you set on Argil subtitles?

The subtitles object has enable (required), an optional styleId, a position and a size. Position and size are three-value enums. That is a small surface: you choose a saved style and nudge it, and the page I read does not list finer controls such as colors or words per line.

What can you set on Sume captions?

Inline avatar captions take enabled, style (default slam), an optional font, a language hint and script_text. The styles are slam, punch, tiktok-green and korean-ad, plus the Hangul identities weight-shift, black-outline, highlight, pill-karaoke, clip-wipe and editorial-emphasis.

For finer control, standalone video captions accept design, which merges over the style: colors (base, active, stroke, accent), typography (font_size_ratio, stroke_width_px and others), placement (anchor_ratio, landscape_anchor_ratio), phrasing (max_words, max_chars, pause_seconds) and motion. Numbers outside range are a 400 at request time instead of a wrong render. design is not supported on punch or tiktok-green.

Comparison

Placement is where the two differ most: Argil gives three positions, Sume gives a numeric line center as a fraction of frame height.

Caption controls, read 2026-10-02
ControlArgilSume
Turn onsubtitles.enablecaptions.enabled, or POST /v1/video-captions
LookstyleIdstyle name, optional design tokens
PositionTop, Middle, Bottomplacement.anchor_ratio (number)
SizeSmall, Medium, Largetypography.font_size_ratio (number)
Existing videoNot on the page readPOST /v1/video-captions with a public video_url

What are the limits on the Sume side?

Inline captions are rejected when the estimated duration is above 60 seconds. A caption failure soft-fails: the avatar job can still succeed with a clean video_url and captions.status=failed, and inline captions do not create a separate billed caption job. A Korean script with a Latin style such as slam is rejected with 400 caption_hangul_text_latin_style, so pick a Hangul style for Korean speech.

Sume burns captions into the picture. It does not export an SRT: the docs say SRT uploads are unsupported and that you pass phrase-level cues or segments instead. If your pipeline needs a sidecar subtitle file, that is a gap. Argil's docs index lists VTT and ASS export, which I did not verify on a page for this post.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume