Blend two videos together API: what Sume can and cannot layer
Sume cannot blend two videos into one frame: video filter works on one clip, and compose puts one still over one video. Two videos join in sequence on Timeline.

Sume has no endpoint that blends two videos into one frame. The video filter takes one clip, and timeline compose layers one still with one video. To combine two videos, place them one after the other as video[] slots on Timeline 1.0, with a transition such as dissolve between them.
Runway's Aug 18, 2026 changelog lists new compositing workflow nodes that blend images and video. This post covers what Sume's docs say about the same job, read 2026-10-01: Timeline compose, Video filter and Timeline 1.0.
Why does the video filter list blend and overlay?
The allowlist in the filter compiler includes split, overlay, hstack, vstack and blend, but the source comment marks them as internal labels only: inputs are still one clip. You can split one clip into two branches and recombine them, for example to build an effect. You cannot feed a second video in; the program has no inputs of its own, and the server wraps the single clip as the source.
What can compose do instead?
Compose takes one still and one video and returns one MP4 with both on screen. operation: "stack" tiles two regions, and operation: "overlay" puts the still on the video with position, width_ratio (0.05-1) and margin_ratio. Fit options are cover, contain, stretch; the docs say blur is not a compose fit.
| Goal | Surface | Note |
|---|---|---|
| Effects on one clip | POST /v1/video-filter | One video_url; allowlisted filtergraph |
| Still over or beside a video | POST /v1/timeline-1.0/compose | operation is stack or overlay |
| Video then video | POST /v1/timeline-1.0/render | Ordered video[], transitions on slots after the first |
| Two videos in one frame | Not offered | Compose requires a still plus a video |
How do I get a dissolve between two videos?
Import both clips with POST /v1/media-imports, then send a Timeline render with two video[] slots and an audio spine (audio.duration_seconds plus audio.url, or audio.mode set to "silence"); the second slot carries transition: { "type": "dissolve", "duration": 0.5 }. Transition duration is capped at 1 second and at 50% of the shorter neighbouring slot. Sizing options are covered in Timeline fit modes.
What if I need a picture-in-picture of two videos?
The docs do not describe one. Compose's image.url must probe as a still, otherwise it returns compose_image_not_still. If you need that layout, produce it outside Sume and import the result. For graphics on a presenter clip, see talking-head video with graphics.
Sources
Related posts
More in Use cases
- Budget for 500 vertical AI video ads a month: $440 to $1,890
500 ten-second vertical clips cost $439 to $1,890 on Sume by Seedance id and resolution; Runway's credit prices give $800 and $1,800 for two of them.
- SB 1050 takedown orders: what an advertising medium must do
SB 1050 bars an advertising medium from running an ad after a served court order. How ad files and URLs fit in, and what a corrected cut means.
- SB 1050 clear and conspicuous disclosure: the definition
SB 1050 defines clear and conspicuous as difficult to miss and easy to read in the ad's medium. Map those words to caption contrast and placement controls.
- California SB 1050: does an AI avatar count as synthetic?
SB 1050 requires a clear and conspicuous disclosure for ads that prominently include a synthetic performer. Read the definition next to avatar creation.
Written by Sume