Cover or contain: fit a 16:9 clip into a vertical stack

Sume timeline compose takes video_fit cover, contain or stretch. Work out what each does to a 16:9 clip in the 1080x1920 default stack, then check one frame.

5 min readSume
All posts

If you stack a still over a 16:9 clip in Sume timeline compose, the fit you choose decides whether viewers see the whole landscape frame or only a centre slice of it. Set layout.video_fit to contain to keep the whole frame inside its region, cover to fill the region and crop the overflow, or stretch to fill it by distorting the picture. Each value is accepted by a stack compose; blur is not a compose fit.

This matters more now that short-form series travel between screens. YouTube's September 23 post says Shorts series is rolling out to creators across web, mobile and TV (YouTube Blog, read 2026-10-04). A shot you frame once gets watched in more than one place, so it is worth making the framing a choice instead of a default.

What the three values are on a stack

Per the timeline compose docs, a stack tiles two regions of one frame. image_fit and video_fit each take cover, contain or stretch. On an overlay compose only video_fit applies. The docs name the three values and do not define the geometry in more detail, so the arithmetic below uses the ordinary meaning of each word. Verify it on your own clip with one extracted frame before you render a batch.

  • cover: scale until the region is full, then crop what spills over.
  • contain: scale until the whole frame fits, leaving part of the region unused.
  • stretch: match the region exactly, ignoring the clip's aspect ratio.

The numbers for 1920x1080 in a 1080x1344 region

Compose defaults to a 1080x1920 output. With ratio: 0.3 on a horizontal split the still takes 576 pixels of height and the video takes the exact remainder, 1344. A 1920x1080 clip therefore has a 1080x1344 region to land in.

Fit arithmetic for a 1920x1080 clip in a 1080x1344 region (worked from the compose docs defaults, read 2026-10-04)
video_fitScale factorSize after scalingWhat the viewer gets
contain0.56251080 x 608Whole frame, region partly empty
cover1.24442389 x 1344, cropped to 1080 wideCentre 45% of the width, no empty space
stretch0.5625 wide, 1.2444 tall1080 x 1344Whole frame, stretched about 2.2x taller than contain

Pick by what is in the frame

A talking head centred in the shot survives cover. A product on the left third of a landscape frame does not. If any detail sits near the left or right edge, use contain and spend the unused region on the still. If you need the full frame and the full region, make the source vertical first with video filter crop instead of letting compose do it.

Compose reads only workspace media.sume.com URLs, so import an outside clip first through POST /v1/media-imports; the media inputs page lists the surface.

Submit it with an explicit fit

This request asks for contain. It waits up to 30 seconds with mode: "sync", answering 200 with a finished job or 202 with a queued one. Replace the two artifact URLs with your own.

import os, requests

body = {
    "operation": "stack",
    "image": {"url": "https://media.sume.com/artifacts/artf_demo/banner.png"},
    "video": {"url": "https://media.sume.com/artifacts/artf_demo/talk.mp4"},
    "layout": {
        "split": "horizontal",
        "image_region": "top",
        "ratio": 0.3,
        "video_fit": "contain",
    },
    "output": {"width": 1080, "height": 1920},
    "mode": "sync",
}
resp = requests.post(
    "https://api.sume.com/v1/timeline-1.0/compose",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "compose-fit-contain-001",
    },
    json=body,
    timeout=60,
)
print(resp.status_code, resp.json().get("request_id"))

Compose is flat $0.02 per job per the docs; confirm the live number in GET /v1/catalog. Change only video_fit between two runs and compare one frame from each with video frames.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume