Product still 30%, clip 70%: compose stack ratio in pixels
Sume compose stack ratio is the still's share of the frame, 0.1 to 0.9, and the video takes the rest. Pixel table for 1080x1920 plus a request that sets 0.3.

In a Sume stack compose, layout.ratio is the still's share of the frame, from 0.1 to 0.9, and the video gets the exact remainder. On the default 1080x1920 output, ratio: 0.3 gives a 576-pixel product still above a 1344-pixel clip. The default 0.5 is the half-banner split.
The reason to care this week: product trends move quickly. Exploding Topics' September 21 update lists massage comb at +12,800%, egg weights at +4,300% and wireless meat thermometer at +2,500% (Exploding Topics, read 2026-10-04). A clip of someone using a gadget with the product photo pinned above it is a common way to show one; the ratio decides how much of the screen the photo gets.
The ratio table for 1080x1920
On a horizontal split the ratio applies to the 1920-pixel height. The still's region is ratio times the height and the video region is what is left. Both numbers below are straight multiplication from the timeline compose docs defaults; they do not include any encoder rounding.
| ratio | Still height | Video height | Typical use |
|---|---|---|---|
| 0.1 (minimum) | 192 px | 1728 px | A thin brand strip |
| 0.2 | 384 px | 1536 px | Price tag over the clip |
| 0.3 | 576 px | 1344 px | Product photo, clip dominant |
| 0.4 | 768 px | 1152 px | Photo and demo about equal |
| 0.5 (default) | 960 px | 960 px | Half-banner |
Vertical splits use width
With split: "vertical" the still takes left or right and the ratio applies to the 1080-pixel width. A 0.4 ratio leaves 432 pixels for the still and 648 for the video, which is narrow for a landscape clip. Keep image_region on the matching axis or the request fails with compose_image_region_wrong_axis; top and bottom belong to a horizontal split.
Request with ratio 0.3
Both URLs must already be this workspace's media.sume.com artifacts. The still must probe as a still and the video as a video. Output length always comes from the video layer, and the still is held for the whole clip.
import os, requests
HEIGHT = 1920
for ratio in (0.2, 0.3, 0.4):
still = round(HEIGHT * ratio)
print(ratio, still, HEIGHT - still)
resp = requests.post(
"https://api.sume.com/v1/timeline-1.0/compose",
headers={
"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "compose-ratio-0-3-001",
},
json={
"operation": "stack",
"image": {"url": "https://media.sume.com/artifacts/artf_demo/product.png"},
"video": {"url": "https://media.sume.com/artifacts/artf_demo/demo.mp4"},
"layout": {"split": "horizontal", "image_region": "top", "ratio": 0.3},
"output": {"width": 1080, "height": 1920},
},
timeout=60,
)
print(resp.status_code, resp.json().get("request_id"))
Compose is flat $0.02 per job; confirm in GET /v1/catalog. Poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result for video_url. Nothing here predicts which ratio performs better; test two ratios on the same clip.
Sources
Related posts
More in Media tools
- Pull a voice track from AI video, then join takes gaplessly
Extract a wav or mp3 from a Sume-hosted video with POST /v1/audio-detach, then join up to 20 audio parts with Timeline audio. Fields, caps, $0.01 rate.
- Swapped the voice? Re-time the visuals from word timestamps
A new voice speaks at a new pace, so cuts set for the old one drift. Take each scene's start from the new take's word timestamps and re-plan the timeline.
- Reference ingest coverage: how many frames OCR read
The reference-ingest coverage block reports frames decoded, frames OCR read and the OCR rate. Use it to decide when on-screen text needs a second look.
- Reference ingest purpose: qa or remix decides who transcribes
Reference ingest purpose defaults to reference_remix, which transcribes speech. brief_format, face_swap and qa do not. An explicit allow_billed_stt wins.
Written by Sume