Storyboard animatic from stills with a Sume timeline
Check pacing and shot order before paying for video: render storyboard stills as a silent Timeline 1.0 animatic, with an unbilled plan call first.

You can build a storyboard animatic on Sume by putting your approved stills into Timeline 1.0 as slots, setting audio.mode to "silence", and rendering one MP4 that shows shot order and timing. Timeline's public rate is $0.10 per output minute (rounded up), so a 60-second animatic costs $0.10, a fraction of a single generated clip.
The point is to decide shot order and length before you spend on image-to-video. If shot 4 should be 2 seconds instead of 5, you find out from a render that costs ten cents, not from eight finished clips.
What does a Sume animatic actually show?
Each still becomes a static hold. The docs say stills are static holds and that a motion field is accepted and ignored, with a motion_ignored warning. So the animatic has no camera moves and no fake parallax. It shows what is on screen, in which order, for how long, and with which cut or fade between shots.
That is less than a previs tool gives you, and it is worth being clear about. If you need to judge camera movement, an animatic from stills will not tell you; a short draft clip will. What the animatic does settle is rhythm: whether the opening shot holds too long, whether a product reveal lands before the voice line, and whether the total fits the platform limit.
How do I set it up without a voiceover?
A render needs an audio spine unless you pass audio.mode: "silence". Silence mode declares a length with no file, and then url, parts, gain_db and source_in are all refused. That is exactly what you want while the script is still moving.
Every source URL must already be a media.sume.com artifact or asset of your workspace. Stills you generated with Sume already are. For a hand-drawn board or a photo from elsewhere, import it first with POST /v1/media-imports; off-host URLs are rejected at admit.
Run the plan call first. POST /v1/timeline-1.0/plan is unbilled, needs no Idempotency-Key, and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros without creating a job.
curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"audio": { "mode": "silence", "duration_seconds": 12 },
"output": { "width": 1920, "height": 1080 },
"video": [
{ "source_url": "https://media.sume.com/artifacts/artf_demo/board-01.png",
"start": 0, "duration": 4, "fit": "contain" },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/board-02.png",
"start": 4, "duration": 3, "fit": "contain",
"transition": { "type": "fade", "duration": 0.25 } },
{ "source_url": "https://media.sume.com/artifacts/artf_demo/board-03.png",
"start": 7, "duration": 5, "fit": "contain" }
]
}'Which settings matter for storyboard frames?
Mixed-ratio boards are common: one panel drawn 4:3, another 16:9. The fit field is set per slot and takes cover (default), contain, stretch or blur. For an animatic, contain is the sensible first try because you want to see the whole panel; render a test and switch if the framing is not what you want.
| Setting | What to use | Why |
|---|---|---|
| audio.mode | silence | Declared length, no spine file needed |
| video[].fit | contain, or cover for full-bleed frames | Default is cover; contain keeps the whole board panel visible |
| output.width / height | Even integers, 256 to 2160 | Default output is 1080x1920, so set 1920x1080 for a landscape board |
| video[].start | 0, then strictly increasing | video[0].start must be 0; declared starts are authoritative |
| video[].duration | 0.2 s or more | Coverage may trail the spine by at most 0.5 s |
| Slots | 1 to 200 | Audio length is 1 to 1800 seconds |
What does the animatic not tell you?
It cannot predict how a model will move a character, and it cannot preview audio. The plan call also cannot predict padding or looping warnings, which only matter for video sources, not stills.
When the order and lengths look right, the same shot list becomes your generation brief. Each still is the first_frame in a frame_images entry on POST /v1/videos, and each slot's duration is the clip length you request. Later you swap the stills for the finished clips in the same video[] array and render again. Because starts are declared rather than derived, you only edit the source_url and duration values.
One cost note: render is $0.10 per rounded-up output minute with no provider inference, so iterating the animatic five times on a 45-second board is still $0.50 in total. Confirm the live rate in GET /v1/catalog before you budget.
Sources
Related posts
More in Use cases
- AI Act marking exemption for B2B and industrial output: how narrow
The Commission FAQ says a narrow Article 50(2) marking exemption is envisaged for B2B or industrial outputs, with conditions in the guidelines. What it lists.
- AI Act Article 50 fines: up to EUR 15M or 3%, and who enforces it
The Commission's Article 50 FAQ says fines can reach 15 million euros or 3% of worldwide turnover, enforced mainly by national market surveillance authorities.
- AI Act Article 50 for non-EU providers: output used in the EU
The Commission FAQ says providers outside the EU are subject to the AI Act if their system's output is used in the EU. How it defines provider and deployer.
- Can one avatar UGC ad change location mid-video? One scene only
Sume's avatar video renders one avatar and one shared scene per final video. For a second location, make two jobs and join them with a Timeline 1.0 render.
Written by Sume