Candle holiday ad with AI: a jar photo as IMAGE_REF_0 and a Lyria bed
Make a candle holiday ad: Gemini Omni Flash reference-to-video from your jar photo, a Lyria 3.5 bed and a Timeline render, about $1.48 on Sume.

To make a candle brand holiday ad with AI, give Gemini Omni Flash 1.1 your jar photo as a reference image and name it in the prompt as <IMAGE_REF_0>. That is reference-to-video: the model builds a scene around the jar you supplied instead of inventing one. A 10-second 720p clip is $1.25, a Lyria 3.5 bed is $0.125 and one Timeline render is $0.10, so $1.475 in total on Sume's public rates.
Candles are a hard product for video models, because the thing a buyer cares about (the flame, the glass, the label) is small. A reference photo is the cheapest insurance, and a bed under the clip does the emotional work that a 10-second silent loop cannot.
What the three jobs cost
Omni is 3 to 10 seconds, 360p to 4K, 16:9 or 9:16, with native synced audio always on. Reference images go up to 10 per request and reference videos up to 3, each at most 3 seconds, and audio references are not accepted. Sume bills Omni at the $0.10 a second 720p list times 1.25, Lyria 3.5 at the $0.10 per generation list times 1.25, and Timeline at $0.10 per output minute rounded up.
| Step | Endpoint | Sume price |
|---|---|---|
| Reference-to-video, 10 s at 720p | POST /v1/video-router/generate | $1.25 |
| Music bed | POST /v1/music-router/generate, sume/music-auto | $0.125 |
| Join clip and bed | POST /v1/timeline-1.0/render | $0.10 |
| Total | $1.475 |
Prompting with a reference
Refer to media by its position in the list you send: the first image is <IMAGE_REF_0>, the second <IMAGE_REF_1>, and videos are <VIDEO_REF_0> and up. A prompt such as "<IMAGE_REF_0> stands on a dark oak shelf at dusk, the flame steady, a slow push-in, soft snow behind the window" keeps the jar as the subject. Say nothing about scent, burn time or wax type, because the footage cannot show them and a claim you cannot support is a listing problem.
If the first render alters the label, send the label crop as a second reference and say that <IMAGE_REF_1> is the exact label. You are limited to 10 image references, so a front, a side and a label crop fit easily.
Keep the first render as a draft. At 360p Omni is $0.03 a second on the list ($0.0375 with the margin), so a 10-second draft is $0.375 against $1.25 at 720p, and you can fix the framing at a third of the price before the 720p render.
Reference videos are limited to 3 seconds each, so a 3-second wick-lighting clip from your own camera can be sent as <VIDEO_REF_0> when you want the flame to look like your flame.
A bed that fits the length
Lyria 3.5 has no duration field, and the router rejects duration and duration_seconds. You control length in the prompt: write "a 12-second track" or time the sections, for example [0:00-0:04] Intro: soft bells. Put exclusions in the positive prompt too, since a non-empty negative_prompt is not supported. The Timeline then trims the bed to the spine with soundtrack.fade_out_seconds (up to 10) and a gain you set in decibels.
- Ask for the bed 2 seconds longer than the clip so the fade has material.
soundtrack.looprepeats a short bed if the track is shorter than the render.duck_db(0 to 20) needs a real audio spine, not silence; refused otherwise withduck_requires_audio_spine.
The bed request
import os
import uuid
import requests
H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": f"candle-bed-{uuid.uuid4()}"}
body = {
"model": "sume/music-auto",
"prompt": ("Warm instrumental holiday bed, soft felt piano and sleigh bells, "
"72 BPM, a 12-second track with a gentle ending. No vocals."),
}
r = requests.post("https://api.sume.com/v1/music-router/generate",
headers=H, json=body, timeout=60)
print(r.status_code, r.json().get("next_action"))
Wire it together
job.model echoes sume/music-auto, and job.request.routed_model names the engine that actually ran, Lyria 3.5 today. Read the audio from result.artifacts[] where type is audio, then pass that URL as the Timeline soundtrack.url. Because Omni already produced a synced audio track, detach it first if you want a voice or sound design under the bed: audio detach is a flat $0.01 and gives you a wav or mp3 to use as audio.url.
Hold the clip's own audio decision until you have heard it. Omni cannot mute its audio (the API rejects generate_audio: false), so a candle clip can arrive with sounds you did not ask for, and the Timeline spine is where you replace them.
Check the order of operations once before you roll it across a range of scents: reference clip, bed, plan, render. A plan call is free, and the full recipe repeats cleanly for every SKU. For the single-reference mechanics see adding yourself to an Omni clip, and for ducking under a voice see the three-minute Short build.
Sources
Related posts
More in Use cases
- Car walkaround clip: a 30-second Seedance 2.5 take for dealer ads
A one-take car walkaround prompt for Seedance 2.5: a 30-second path in 16:9, the Sume price at 720p and 1080p, and what to verify before posting.
- Cart-abandoner retargeting cuts from one hero ad: three trims, $0.06
TikTok's holiday guide says to retarget cart abandoners and video engagers at peak. Cut a 15-second hero into three retargeting clips with video-trim for $0.06.
- Change 2026 to 2027 on 30 old graphics: an Ideogram 4.5 edit batch
Refresh the year on 30 finished graphics with Ideogram 4.5 edits on Sume: 30 calls cost $1.13 at low, what the queue accepts per plan, and what to proofread.
- Change one label in an infographic image: an Ideogram 4.5 edit
Fix one word or number in a finished infographic PNG with a single edit call on Sume. The prompt names the label, its position and what stays put.
Written by Sume