Stream starting soon screen: make an animated loop
A stream starting soon screen is a looping video the size of your stream's canvas. How to make an animated one with AI: the loop, the words, the length.

A stream starting soon screen is a short looping video, or a still image, that you show full-screen before you go live, with words such as “Starting soon” and your channel name. Make it the size of your stream's canvas (1920×1080 for a 1080p stream, 1280×720 for 720p), and hide the loop point by ending the clip on the same frame it starts with.
With Sume that is three steps: an image-to-video clip that starts and ends on your background image, one caption job that burns the words, and a Timeline render that repeats the clip for as long as you need. Facts come from the Video generation, Video captions, and Timeline 1.0 docs and the Sume API reference, read on 2026-09-28; anything called current code is read from Sume's source.
What size should a starting soon screen be?
The same size as your stream's output, so it fills the frame without being scaled: 1920×1080 for a 16:9 stream at 1080p, 1280×720 at 720p. Timeline 1.0 renders any even size from 256 to 2,160 pixels per edge, so both work, but a 2560×1440 canvas is wider than that limit.
How do I make an animated starting soon screen with AI?
Start from a 16:9 background image at a public HTTPS URL, your own art or one you generated. Leave the words out of the picture: they go on in the next step, as text you write.
- Send the image to
POST /v1/videosas both thefirst_frameand thelast_frame, and describe calm motion that can come back to where it started. In the current catalog, every listed model takes a last frame exceptgrok-imagine-video-1.5; Seamless loop AI video shows how to check the loop. - Keep the clip's sound on;
generate_audiodefaults to what the model can do. In current code the caption job refuses a clip with no audio stream, and the Timeline step drops the clip's own audio anyway. - When the job completes, read the clip's
media.sume.comURL fromGET /v1/jobs/{id}/result.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: starting-soon-loop-v1" \
-d '{
"model": "wan-3.0",
"prompt": "Slow drifting clouds and soft neon light over a city skyline; the camera stays still",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/starting-soon-bg.png" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/starting-soon-bg.png" }, "frame_type": "last_frame" }
],
"aspect_ratio": "16:9",
"resolution": "1080p",
"duration": 10
}'How do I add the “Starting soon” text?
Burn it with POST /v1/video-captions: one cue whose text is your line, with start at 0 and end at the clip's length. Cues skip speech-to-text and burn your words at those times; How to add text over a video covers where the line sits.
- In current code the default
slamstyle prints the words in capitals and lays a light black tint over the whole frame. - Put “Starting soon” and your channel name in the same cue: in current code
slamdraws every cue at the same height and ignores line breaks. - Set
design.motion.enter_secondsandexit_secondsto 0 (each takes 0–2) so the words have no entrance or exit animation at the loop point.
How do I make it run as long as my pre-show?
Loop the captioned clip in one Timeline 1.0 slot as long as the screen should run, with render.pad_mode set to loop; How to increase video length by looping explains the mechanics. For a pre-show there is no voice, so set audio.mode: "silence" with the length in audio.duration_seconds, up to 1,800 seconds (30 minutes), and set output to 1920×1080, since the default frame is 1080×1920.
- Timeline reads only your workspace's
media.sume.comfiles. Completed Sume jobs can include public artifacts there, including the captioned clip. - For music, add a Sume-hosted
soundtrackwithloop: trueandgain_dbat 0; the default −16 dB is a bed level under a voice. AI lofi video generator pairs the same kind of loop with a track.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: starting-soon-30min-v1" \
-d '{
"audio": { "mode": "silence", "duration_seconds": 1800 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/starting-soon.mp4", "start": 0, "duration": 1800 }
],
"output": { "width": 1920, "height": 1080 },
"render": { "pad_mode": "loop" }
}'Can the screen show a live countdown?
No. A rendered file can't know when you'll actually go live, so any countdown in it is fixed. For a short one, caption a separate clip with one cue per second and the same zero enter and exit times: a caption job takes up to 200 cues, each timed within 60 seconds, and in current code its source must be 60 seconds or less with an audio stream, such as a 30-second generated clip with its sound on. Play it after the loop, not inside it, or the count restarts on every pass.
What are the limits, and what does it cost?
The clip is reserved at its model's provider list price × 1.25. Timeline is listed at $0.10 per output minute and reserves whole minutes, so a 30-minute screen reserves $3.00. Both are plus a 5.5% agent fee by default. A caption job is a fixed amount per clip of up to 60 seconds, listed on the Video captions page.
| Step | Limit |
|---|---|
| Loop clip length | 2–30 seconds by model: seedance-2.5 and wan-3.0 go to 30, gemini-omni-flash-1.1 stops at 10, and every other model tops out at 15 |
| Last frame | Every listed model except grok-imagine-video-1.5 (current catalog) |
| Caption cues | 1–200 per job; text 1–400 characters; times 0–60 seconds |
| Caption source | 60 seconds or less, with an audio stream (current code) |
| Timeline render | 1–1,800 seconds; each edge an even 256–2,160 pixels |
Sources
Related posts
More in Use cases
- Text to speech time calculator: estimate, then measure
Estimate speech time as word count divided by a speaking rate, then measure the real file: Sume's TTS reports its duration with word timestamps.
- TikTok Commercial Music Library (CML): what it covers
TikTok's Commercial Music Library is a pre-cleared set of songs businesses may use free, but only on TikTok. What it covers, and music for other uses.
- Transcribe a lecture to notes with timestamps
Transcribe a lecture with sentence timestamps, then have a language model turn it into notes whose headings point back to where each topic starts.
- Video prospecting with an AI avatar: one clip per prospect
Video prospecting puts a short personal video in a sales outreach message. With an AI avatar, fill one script template per prospect and render each clip.
Written by Sume