Storyboard animatic from stills with a Sume timeline

Check pacing and shot order before paying for video: render storyboard stills as a silent Timeline 1.0 animatic, with an unbilled plan call first.

5 min readSume
All posts

You can build a storyboard animatic on Sume by putting your approved stills into Timeline 1.0 as slots, setting audio.mode to "silence", and rendering one MP4 that shows shot order and timing. Timeline's public rate is $0.10 per output minute (rounded up), so a 60-second animatic costs $0.10, a fraction of a single generated clip.

The point is to decide shot order and length before you spend on image-to-video. If shot 4 should be 2 seconds instead of 5, you find out from a render that costs ten cents, not from eight finished clips.

What does a Sume animatic actually show?

Each still becomes a static hold. The docs say stills are static holds and that a motion field is accepted and ignored, with a motion_ignored warning. So the animatic has no camera moves and no fake parallax. It shows what is on screen, in which order, for how long, and with which cut or fade between shots.

That is less than a previs tool gives you, and it is worth being clear about. If you need to judge camera movement, an animatic from stills will not tell you; a short draft clip will. What the animatic does settle is rhythm: whether the opening shot holds too long, whether a product reveal lands before the voice line, and whether the total fits the platform limit.

How do I set it up without a voiceover?

A render needs an audio spine unless you pass audio.mode: "silence". Silence mode declares a length with no file, and then url, parts, gain_db and source_in are all refused. That is exactly what you want while the script is still moving.

Every source URL must already be a media.sume.com artifact or asset of your workspace. Stills you generated with Sume already are. For a hand-drawn board or a photo from elsewhere, import it first with POST /v1/media-imports; off-host URLs are rejected at admit.

Run the plan call first. POST /v1/timeline-1.0/plan is unbilled, needs no Idempotency-Key, and returns duration_seconds, segment_count, billable_minutes and estimated_cost_usd_micros without creating a job.

curl -X POST https://api.sume.com/v1/timeline-1.0/plan \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio": { "mode": "silence", "duration_seconds": 12 },
    "output": { "width": 1920, "height": 1080 },
    "video": [
      { "source_url": "https://media.sume.com/artifacts/artf_demo/board-01.png",
        "start": 0, "duration": 4, "fit": "contain" },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/board-02.png",
        "start": 4, "duration": 3, "fit": "contain",
        "transition": { "type": "fade", "duration": 0.25 } },
      { "source_url": "https://media.sume.com/artifacts/artf_demo/board-03.png",
        "start": 7, "duration": 5, "fit": "contain" }
    ]
  }'

Which settings matter for storyboard frames?

Mixed-ratio boards are common: one panel drawn 4:3, another 16:9. The fit field is set per slot and takes cover (default), contain, stretch or blur. For an animatic, contain is the sensible first try because you want to see the whole panel; render a test and switch if the framing is not what you want.

Timeline settings that matter for an animatic (Sume docs, read 2026-10-02)
SettingWhat to useWhy
audio.modesilenceDeclared length, no spine file needed
video[].fitcontain, or cover for full-bleed framesDefault is cover; contain keeps the whole board panel visible
output.width / heightEven integers, 256 to 2160Default output is 1080x1920, so set 1920x1080 for a landscape board
video[].start0, then strictly increasingvideo[0].start must be 0; declared starts are authoritative
video[].duration0.2 s or moreCoverage may trail the spine by at most 0.5 s
Slots1 to 200Audio length is 1 to 1800 seconds

What does the animatic not tell you?

It cannot predict how a model will move a character, and it cannot preview audio. The plan call also cannot predict padding or looping warnings, which only matter for video sources, not stills.

When the order and lengths look right, the same shot list becomes your generation brief. Each still is the first_frame in a frame_images entry on POST /v1/videos, and each slot's duration is the clip length you request. Later you swap the stills for the finished clips in the same video[] array and render again. Because starts are declared rather than derived, you only edit the source_url and duration values.

One cost note: render is $0.10 per rounded-up output minute with no provider inference, so iterating the animatic five times on a 45-second board is still $0.50 in total. Confirm the live rate in GET /v1/catalog before you budget.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume