How to combine photos and videos into one video

Combine photos and videos into one video on a single timeline: each photo held for a set time, each clip cut to length, one soundtrack under it all.

5 min readSume
All posts

To combine photos and videos into one video, lay them out on one timeline in the order you want, give each photo a length on screen, cut each clip to the part you want, pick one frame size for everything, and put one soundtrack under the whole sequence. With Sume, that is one Timeline 1.0 render: every photo and clip is a video[] slot, a photo is held still for its slot's duration, and a clip plays from its source_in.

The facts come from the Timeline 1.0, Audio detach, and Video frames docs and the Sume API reference, read on 2026-09-28. Anything described as current behavior is read from Sume's code. A slideshow of photos alone is covered in turn Sume images into an MP4 with music.

How do I put photos and clips in one render?

Every URL must already be your workspace's media.sume.com artifact or asset, such as an earlier Sume job's output: a generated clip, or a still pulled from a clip with video frames, which returns durable media.sume.com images. There is no public upload route for a file on your computer; which URLs each endpoint accepts explains the rule.

Each slot starts where the previous one ends, and audio.duration_seconds is where the last one stops (11 seconds here). This render opens on a photo, fades to five seconds of a clip starting two seconds in, then fades to a second photo, with music under all of it:

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: photos-and-clips-001" \
  -d '{
    "output": { "width": 1920, "height": 1080 },
    "audio": { "mode": "silence", "duration_seconds": 11 },
    "soundtrack": { "url": "https://media.sume.com/artifacts/artf_demo/music.mp3", "gain_db": 0, "loop": true },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/photo-1.png", "start": 0, "duration": 3, "fit": "blur" },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/clip.mp4", "start": 3, "duration": 5, "source_in": 2,
        "transition": { "type": "fade", "duration": 0.5 } },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/photo-2.png", "start": 8, "duration": 3, "fit": "blur",
        "transition": { "type": "fade", "duration": 0.5 } }
    ]
  }'

What changes between a photo slot and a clip slot?

Both are video[] slots with the same fields, but a photo has no running time, no frame rate, and no sound of its own, so the render handles the two differently:

From Timeline 1.0 and the Sume API reference, read 2026-09-28; cells marked current code are read from Sume's compiler.
WhatPhoto slotClip slot
On screenFor the slot's duration (at least 0.2 s), held stillFor the slot's duration (at least 0.2 s), playing from source_in
source_inIgnored, with a warning (current code)The in-point into the file
Frame rateNoneWith output.fps omitted, the longest video sources set the output rate
Its own soundNoneDropped unless you put it on the audio spine (current code)
Source shorter than the slotNever happens: a still covers its whole slot (current code)Loops or holds its last frame, as render.pad_mode says
A transition into the slotAlways has a picture to fade to (current code)Needs source_in + duration + the transition's length of footage, or the boundary becomes a hard cut (current code)

How do I mix portrait and landscape shots?

Pick one frame with output. It is 1080×1920, a vertical frame, when omitted, so set 1920×1080 for a landscape video, as the example does. Then give each slot a fit: in current code the default, cover, fills the frame and crops what overflows, which can cut the top and bottom off a portrait photo. How to merge two videos of different resolutions compares the four fits.

Where does the sound come from?

From the audio spine and the optional soundtrack only. In current code a render maps no clip's own audio, so a clip's sound is dropped unless you put it on the spine.

  • Music only: audio.mode: "silence" sets the length and the music plays as the soundtrack, as in the example; the photos-only guide explains the settings that track needs.
  • A voiceover: pass it as audio.url and lay the music under it as a soundtrack bed, as in add background music to a video.
  • A clip's own sound: detach it with POST /v1/audio-detach and use the file as one of up to 20 audio.parts[], trimmed to the stretch its slot plays; remove part of a video by API shows the pattern. The parts must together cover audio.duration_seconds, so the photo slots need a part as well, such as a slice of the music. In current code all parts must share one channel layout, or the render fails with audio_parts_channel_mismatch.
  • Every audio URL, the soundtrack included, must be Sume-hosted too.

What does it cost, and what are the limits?

A render is listed at $0.10 per output minute on API pricing, plus a 5.5% agent fee by default, and reserves ceil(audio.duration_seconds / 60) minutes. POST /v1/timeline-1.0/plan checks the same body without billing. One render holds 1 to 200 slots and 1 to 1,800 seconds of output. A transition can go on any slot after the first, photo or clip, and lasts at most 1 second and half the shorter neighboring slot; the types are compared in video transitions API.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume