Gift guide video from 20 product stills: Timeline plan and render
Hold 20 product stills for 3 seconds each over one voiceover in Timeline 1.0. Plan it free, expect a chunked render past 12 slots, and cap fades at 8 in a row.

To make a gift guide from 20 product stills, put each still in a video[] slot of Sume's Timeline 1.0, hold it for 3 seconds, and lay one voiceover file on the audio spine. Twenty 3-second slots make a 60-second 1080x1920 MP4, which at $0.10 per output minute is $0.10. Run POST /v1/timeline-1.0/plan first; it is unbilled and returns the duration, segment count and estimated cost without creating a job.
A still is a static hold. This is an assembly tool, not an animation model, so a gift guide built this way is a slideshow with a voiceover. If you want motion on each product, generate short clips first and place those in the slots instead.
Why does 20 slots change how the render runs?
render.strategy defaults to auto, which chunks the render once there are more than 12 segments. Forcing single with more than 12 slots is refused with render_strategy_unsafe. So a 20-product guide renders chunked unless you ask otherwise, and you should leave the default alone.
The limits for the document itself are generous: 1 to 200 video slots and an audio spine of 1 to 1800 seconds. Every URL must already be a media.sume.com artifact or asset in your workspace, so import product photos first, as described in Media inputs.
How do transitions behave across 20 slots?
Transitions live on slots after the first. The type is one of fade, wipeleft, wiperight, slideup, slidedown or dissolve, the duration is at most 1 second, and it cannot exceed 50% of the shorter neighbouring slot. More than 8 adjacent fades is refused with too_many_chained_transitions; the docs say to insert a hard cut.
The script below puts a 0.2-second fade on every second slot and sets output.fps to 30, so 0.2 seconds is a whole six frames (the docs refuse transitions that are not frame-aligned), so no run of fades gets near the limit and every other cut stays hard. Declared start values are authoritative: the compiler compensates for the crossfade, it never shifts your starts.
import json, os, urllib.request
KEY = os.environ["SUME_API_KEY"]
BASE = "https://media.sume.com/artifacts/demo/"
video = []
for i in range(20):
slot = {"source_url": f"{BASE}gift-{i + 1:02d}.jpg",
"start": i * 3, "duration": 3, "fit": "cover"}
if i % 2 == 1:
slot["transition"] = {"type": "fade", "duration": 0.2}
video.append(slot)
doc = {"audio": {"url": BASE + "voice.wav", "duration_seconds": 60},
"soundtrack": {"url": BASE + "bed.mp3", "gain_db": -6, "duck_db": 8},
"output": {"fps": 30}, "video": video}
req = urllib.request.Request(
"https://api.sume.com/v1/timeline-1.0/plan", json.dumps(doc).encode(),
{"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
print(json.load(r))What do you read from the plan before paying?
The plan response gives duration_seconds, segment_count, billable_minutes, estimated_cost_usd_micros and a filtergraph_summary. It runs the schema checks and the Sume-host URL checks, so a typo in one of the 20 image URLs shows up here instead of in a failed render. It does not download media and cannot predict padding or looping warnings for short sources, which is not a concern for stills.
When the plan looks right, send the same document to POST /v1/timeline-1.0/render with an Idempotency-Key. The default mode is async, so you poll GET /v1/jobs/:id/status and read GET /v1/jobs/:id/result, which returns video_url, duration_seconds and segment_count.
| Setting | Value used | Rule from the docs |
|---|---|---|
| Slots | 20 | 1 to 200 per timeline |
| Slot length | 3 s | Minimum 0.2 s; coverage may trail the spine by at most 0.5 s |
| Spine | 60 s voiceover | 1 to 1800 s |
| Output | 1080x1920 default | Even integers 256 to 2160 if overridden |
| Fades | Every second slot, 0.2 s at 30 fps | At most 8 adjacent; at most 1 s |
| Music | Bed at -6 dB, duck 8 dB | duck_db 0 to 20; needs a real spine, not silence |
| Price | $0.10 | Per ceil output minute |
What does this not give you?
There is no price text, product name or logo drawn by the timeline. If you need burned-in overlay copy, add it with video captions using authored cues, which skip speech-to-text. If a product photo has a different aspect ratio from the output, fit decides what happens: cover crops by default, contain letterboxes, and blur fills the sides with a blurred copy.
Sources
Related posts
More in Use cases
- GPT Image 2.5 ad headline text: quote the copy, test medium vs high
For an ad still with a headline, quote the exact words, state position and type style, and render at medium and high to compare. A Sume script and its cost.
- Ad still from a packshot: tell GPT Image 2.5 what must not move
Turn a product packshot into an ad still with GPT Image 2.5 by listing what must not move: shape, colour, label text. A Sume request and a checklist.
- GPT Image 2.5 drafts at low quality, then a final: cost and steps
fal lists GPT Image 2.5 at $0.00588 low and $0.05268 high for 1024 squares. Draft four options at low, pick one, then redo it at high. Sume steps and the catch.
- Keep the same face in a GPT Image 2.5 edit: the prompt block to use
To keep a person recognisable in a GPT Image 2.5 edit, send the photo as a reference and spell out what must not change. The exact wording and a Sume call.
Written by Sume