AI whiteboard animation generator: blank board to sketch
AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.

An AI whiteboard animation generator can imitate the hand-drawn explainer look: give a video model a blank-board still as the first frame and the finished sketch as the last frame, and prompt the drawing in between, one idea per clip. Generated lettering can come out wrong, so draw the pictures with AI and burn the words on as captions, with the narration setting the pace.
Facts come from Sume's Video generation, Image API, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. The strokes and any drawing hand are generated, so they may not move like real drawing; watch each clip.
How does AI make a whiteboard drawing animation?
With two stills per idea. The first is the board before the idea is drawn; the last is the board after. A video model fills the frames between them, and the prompt describes the drawing, such as “a marker sketches a lightbulb from left to right, still camera, plain white board”.
- Make the finished sketch from the blank board: send the board still as a reference in
input_referencesonPOST /v1/imagesand ask for the line drawing on it, so both frames share one board. References must be public HTTPS, and models whoseinput_referencesdescriptor is{"min": 0, "max": 0}reject them. The Image API returns the sketch at a signed Sume URL, so copy the file you keep and host it at your own public HTTPS URL before you use it as a frame. - For the next idea, the last sketch becomes the next clip's first frame, so the board fills up from shot to shot.
- Only models whose
supported_frame_imageslistslast_frametake an end frame, and in current code alast_framewithout afirst_frameis refused.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: board-bulb-001" \
-d '{
"model": "seedance-2",
"prompt": "A black marker sketches a lightbulb on a white board, still camera",
"frame_images": [
{ "type": "image_url", "image_url": { "url": "https://example.com/board-blank.png" }, "frame_type": "first_frame" },
{ "type": "image_url", "image_url": { "url": "https://example.com/board-bulb.png" }, "frame_type": "last_frame" }
],
"resolution": "720p",
"aspect_ratio": "16:9",
"duration": 6
}'Can the AI write the words on the whiteboard?
Don't count on it. A model may misspell a label or garble a number, and fixing one word means generating the clip again. Keep text out of the drawings and put it on with a caption job: each cue burns exactly the text you send between its start and end seconds. Left without a style, Latin text gets slam, which in current code sets words in capitals and lays a light dark overlay over the whole frame, dimming the white board; set a style from the Video captions page if that matters.
In current code the caption job refuses a video over 60 seconds or one without an audio stream, so caption the finished, narrated video, in parts under a minute if it runs longer. Add captions to a long video shows the split.
How do I turn the clips into a whiteboard video?
Write one sentence per idea and voice the script with POST /v1/tts-1.0/generate. With timestamps.words: true and segmentation.mode: "sentence", the result carries sentence segments[] with start and end times, so each drawing clip can start where its sentence starts. Slideshow with AI voiceover shows that mapping.
Then one POST /v1/timeline-1.0/render puts the narration on the audio spine and the clips in video[] at those starts. When a sentence outlasts its clip, render.pad_mode: "freeze" holds the last frame, the finished sketch, instead of replaying the drawing. In current code each clip's own sound is dropped, so viewers hear the narration and an optional soundtrack.
How much does an AI whiteboard animation cost?
A clip per idea, one narration, one render and one caption job per minute of video. Sketch stills from the Image API are priced per model, listed on GET /v1/images/models.
| Step | Call | Price |
|---|---|---|
| Each drawing clip | POST /v1/videos | By model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models |
| Narration | POST /v1/tts-1.0/generate | $0.0475 per 1,000 characters |
| Join | POST /v1/timeline-1.0/render | $0.10 per output minute, reserved in whole minutes |
| Words on screen | POST /v1/video-captions | $0.20 per job, for videos up to 60 seconds |
What are the limits?
- No video model accepts a
seed, so a redrawn clip comes out different; keep the takes you approve. - Stills for generation must be at public HTTPS URLs; the render takes only this workspace's
media.sume.comfiles, such as the clips Sume returned. - The render's default frame is 1080×1920, vertical; set
output.widthandoutput.heightto match 16:9 clips. - For the presenter or B-roll style of explainer instead, see How to make an explainer video with AI.
Sources
Related posts
More in Use cases
- Background music for commercials: generate a bed that fits
Background music for a commercial is an instrumental bed under the voiceover that ends on time. How to brief, generate, cut and duck one with AI.
- Can I use AI images on Amazon KDP? What you must disclose
Yes, but KDP requires you to disclose AI-generated images, including cover and interior art. AI-assisted edits of your own work need no disclosure.
- Can I use AI images on Pinterest? What the AI label means
Yes. Pinterest labels Pins it detects as AI-generated or AI-edited "AI modified", from metadata or its classifiers, and you can appeal a wrong label.
- Car background removal AI: one backdrop for every vehicle
Car background removal AI cuts each car out as a transparent PNG at a flat per-image price, or redraws the setting while the car stays. Steps and cost.
Written by Sume