WebVTT cue settings (line, size) vs Sume anchor_ratio and width

WebVTT line:78%,center and size:90% have close matches in Sume's caption design fields. A converter script, and what per-cue settings a burned render loses.

5 min readSume
All posts

WebVTT lets each cue carry its own line, position, size and align settings after the timing arrow, and a percentage line value is an offset from the top of the video viewport (read 2026-10-03). Sume's burned-in captions have no per-cue settings, but two request-level fields correspond closely: design.placement.anchor_ratio for vertical position and design.typography.safe_width_ratio for width. Everything else in a WebVTT cue-settings list has no Sume equivalent.

The script below converts a cue-settings string to a Sume design block and says what it ignores.

What do WebVTT cue settings control?

In the W3C WebVTT specification, the settings that may follow a cue's timing are vertical, line, position, size, align and region, and each may appear once per cue. For horizontal text, line sets the vertical offset of the cue box from the top of the viewport, either as a line number or as a percentage, with an optional start, center or end alignment, and start is the default. position is a percentage across the width for horizontal cues, and the spec's sample file uses values such as size:35% and align:right. Cues using vertical, line or size cannot belong to a region (read 2026-10-03).

Sume documents anchor_ratio as the caption line centre as a fraction of frame height, so a WebVTT line:78%,center is the nearest equivalent, and size:90% is the nearest to safe_width_ratio: 0.9 (video captions).

WebVTT cue settings against Sume design fields, read 2026-10-03
WebVTT settingMeaningSume equivalent
line:78%,centerCue box centre 78% down the viewportplacement.anchor_ratio 0.78 (range 0.05 to 0.95)
line:3 (a line number)Counted in lines, not a fractionNone; convert to a percentage first
size:90%Cue box width as a share of the viewporttypography.safe_width_ratio 0.9 (range 0.3 to 1)
position:, align:Horizontal placement and text alignmentNone documented
vertical:, region:Vertical text, scrolling regionsNone

How do I convert them?

The script maps percentage line and size values and reports the settings it drops. It rejects a line-number line value, because a count of lines depends on the player's font size and cannot be turned into a fraction of the frame without guessing. Remember the fields are request-wide: one anchor_ratio and one width for the whole render, while WebVTT can vary them per cue.

Use it to carry the position of a typical cue over from an existing sidecar when you burn the same clip.

import json, re


def vtt_settings_to_design(settings: str) -> dict:
    """Map WebVTT 'line:78%,center size:90%' to a Sume design override."""
    design = {}
    for part in settings.split():
        name, _, value = part.partition(":")
        if name == "line":
            pct = re.match(r"(\d+(?:\.\d+)?)%(?:,(start|center|end))?$", value)
            if not pct:
                raise ValueError("only percentage line values map: " + value)
            if pct.group(2) not in (None, "center"):
                print("note: line alignment is not centre; check the render")
            design["placement"] = {"anchor_ratio": float(pct.group(1)) / 100}
        elif name == "size":
            ratio = float(value.rstrip("%")) / 100
            design["typography"] = {"safe_width_ratio": ratio}
        elif name in ("position", "align", "vertical", "region"):
            print("ignored, no Sume field:", part)
    return design


body = {
    "video_url": "https://media.sume.com/artifacts/example/vertical.mp4",
    "style": "slam",
    "design": vtt_settings_to_design("line:78%,center size:90% align:center"),
}
print(json.dumps(body, indent=2))

Why does the vertical position need care?

A percentage line in WebVTT is measured from the top edge and the line alignment value says which part of the cue box sits at that offset: the start, centre or end, with start as the default. That default matters. A cue with line:78% and no alignment puts the top of the box at 78%, so a multi-line cue would run well below it. Sume's anchor_ratio is defined by the line centre, which is why the script above expects center and prints a note when it sees anything else. If your file uses the default start, shift the number up by roughly half the height of the caption block before you convert, then check one rendered frame.

The width setting is simpler, since size and safe_width_ratio are both shares of the frame width. Note that WebVTT counts the cue box and Sume counts the width a caption line may occupy, so they match in intent rather than to the pixel.

What is lost in the conversion?

Burning captions fixes the layout for the whole clip, so some WebVTT behaviour cannot carry over.

  • Per-cue moves: a cue that jumps up to avoid a lower-third graphic cannot move in a single Sume render.
  • Left and right alignment for two-speaker dialogue, as in the spec's interview sample.
  • Region scrolling, where up to three lines roll as new cues arrive.
  • Vertical text: it has no Sume setting, so burn it some other way or leave it as a sidecar.
  • Player behaviour: a sidecar can be repositioned or styled by the viewer's player, a burned render cannot.

When should I keep the sidecar instead?

If the destination accepts a WebVTT file, the track keeps every setting above and stays editable, so burn only where a track is not accepted or the clip will be reposted. The sidecar post covers building the VTT from a Sume transcript. For choosing an anchor_ratio for a given app's interface, see the Shorts safe zone post.

A caption job is $0.20 per job for videos up to 60 seconds under the current fixed estimate; confirm live pricing in GET /v1/catalog. Out-of-range design values are rejected with a 400 at request time.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume