16:9 insert in a 1080x1920 Short: Timeline fit modes compared
In a 1080x1920 Timeline, cover keeps 31.6% of a 16:9 clip's width, contain shrinks it to 608 px tall, stretch distorts it and blur pads it.

When a 16:9 clip goes into a vertical Timeline, video[].fit decides how it fills the frame. cover, the default, scales it up and crops, keeping about 31.6% of the width. contain fits the whole picture and pads with black. stretch fills the frame and distorts. blur fits the whole picture over a blurred copy of itself.
What each does at 1080x1920
Behaviour is from the Timeline compiler source on main; the numbers are arithmetic for a 1920x1080 source in a 1080x1920 frame.
| fit | What happens | Picture kept |
|---|---|---|
| cover (default) | Scale up to fill, crop the overflow | About 31.6% of the width, full height |
| contain | Scale down to fit, pad with black | Whole picture, about 608 px of the 1920 px height |
| stretch | Scale to the output size | Whole picture, distorted |
| blur | Contained copy over a blurred, cover-cropped copy | Whole picture, no black bars |
Picking one
Use cover when the subject is centred and the frame is more important than the edges, such as a face. Use contain for a screen recording or a chart where cropping would cut data, and accept the bars. Use blur when you want the whole picture and bars would look unfinished. Avoid stretch except for abstract footage.
A vertical platform page does not tell you which is right. TikTok's ad page recommends 9:16 and lists a 540x960 minimum, and YouTube's page says Shorts are vertical. Neither says anything about bars. That is a design choice for you.
Mixing fits in one Short
Fit is set per slot, not per render, so one Short can mix them. A generated vertical clip needs no fit decision, because it already matches the frame. A landscape screen recording can use contain or blur, and a landscape b-roll shot can use cover.
Be consistent about the rule you choose. Switching between black bars and blur from shot to shot looks accidental, even if each choice made sense alone. Pick one treatment for non-vertical sources and apply it throughout.
Also keep the first frame in mind. In many apps the first frame is what shows before playback, and a bar-heavy frame looks weaker in a feed than a full-frame one. That is a design view, not a spec; neither page I read says anything about it.
The cost
A blur fit builds extra filter nodes per slot, and the compiler downsizes the blur background to an eighth of the frame to keep it cheap. Timeline is billed by output minute, not by fit mode, at the public rate in the doc ($0.10 per output minute, rounded up), so the choice is about looks and render time. See the 4:5 on a 9:16 canvas comparison for another ratio.
{
"audio": { "mode": "silence", "duration_seconds": 6 },
"video": [{ "source_url": "https://media.sume.com/artifacts/artf_demo/wide.mp4", "start": 0, "duration": 6, "fit": "blur" }],
"output": { "width": 1080, "height": 1920 }
}Sources
Related posts
More in Media tools
- 45-minute webinar into YouTube Shorts: Sume trim takes 30 minutes
Sume video trim accepts a source up to 1800 seconds and cuts up to 900. A 45-minute webinar needs a split first; YouTube Shorts run up to 3 minutes.
- 720p landscape to a TikTok 9:16 ad: a 405-wide crop misses 540x960
Cropping 16:9 to 9:16 keeps 31.6% of the width. A 1080p source gives about 608x1080 and clears TikTok's 540x960 minimum; 720p gives 405x720 and does not.
- How to turn an AI image into an og:image link preview in Python
Ideogram 4.5 on Sume makes a 2:1 image at 1K as 1440x720. A Pillow crop to your link-preview size, plus the og:image:alt tag Open Graph asks for.
- Burn in captions from your own word timings with Sume (no STT)
Send words or cues to Sume video captions and it burns your text at your times with no speech-to-text. Send only one text input per request.
Written by Sume