AI whiteboard animation generator: blank board to sketch

AI can imitate whiteboard animation: a blank-board first frame, a finished-sketch last frame, the drawing in between. Words go on as captions.

5 min readSume
All posts

An AI whiteboard animation generator can imitate the hand-drawn explainer look: give a video model a blank-board still as the first frame and the finished sketch as the last frame, and prompt the drawing in between, one idea per clip. Generated lettering can come out wrong, so draw the pictures with AI and burn the words on as captions, with the narration setting the pace.

Facts come from Sume's Video generation, Image API, Timeline 1.0 and Video captions docs and the Sume API reference, read on 2026-09-29. Limits marked as current behavior are read from Sume's code. The strokes and any drawing hand are generated, so they may not move like real drawing; watch each clip.

How does AI make a whiteboard drawing animation?

With two stills per idea. The first is the board before the idea is drawn; the last is the board after. A video model fills the frames between them, and the prompt describes the drawing, such as “a marker sketches a lightbulb from left to right, still camera, plain white board”.

  • Make the finished sketch from the blank board: send the board still as a reference in input_references on POST /v1/images and ask for the line drawing on it, so both frames share one board. References must be public HTTPS, and models whose input_references descriptor is {"min": 0, "max": 0} reject them. The Image API returns the sketch at a signed Sume URL, so copy the file you keep and host it at your own public HTTPS URL before you use it as a frame.
  • For the next idea, the last sketch becomes the next clip's first frame, so the board fills up from shot to shot.
  • Only models whose supported_frame_images lists last_frame take an end frame, and in current code a last_frame without a first_frame is refused.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: board-bulb-001" \
  -d '{
    "model": "seedance-2",
    "prompt": "A black marker sketches a lightbulb on a white board, still camera",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/board-blank.png" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/board-bulb.png" }, "frame_type": "last_frame" }
    ],
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "duration": 6
  }'

Can the AI write the words on the whiteboard?

Don't count on it. A model may misspell a label or garble a number, and fixing one word means generating the clip again. Keep text out of the drawings and put it on with a caption job: each cue burns exactly the text you send between its start and end seconds. Left without a style, Latin text gets slam, which in current code sets words in capitals and lays a light dark overlay over the whole frame, dimming the white board; set a style from the Video captions page if that matters.

In current code the caption job refuses a video over 60 seconds or one without an audio stream, so caption the finished, narrated video, in parts under a minute if it runs longer. Add captions to a long video shows the split.

How do I turn the clips into a whiteboard video?

Write one sentence per idea and voice the script with POST /v1/tts-1.0/generate. With timestamps.words: true and segmentation.mode: "sentence", the result carries sentence segments[] with start and end times, so each drawing clip can start where its sentence starts. Slideshow with AI voiceover shows that mapping.

Then one POST /v1/timeline-1.0/render puts the narration on the audio spine and the clips in video[] at those starts. When a sentence outlasts its clip, render.pad_mode: "freeze" holds the last frame, the finished sketch, instead of replaying the drawing. In current code each clip's own sound is dropped, so viewers hear the narration and an optional soundtrack.

How much does an AI whiteboard animation cost?

A clip per idea, one narration, one render and one caption job per minute of video. Sketch stills from the Image API are priced per model, listed on GET /v1/images/models.

From Video generation, Timeline 1.0, Video captions and the Sume API reference, read 2026-09-29. Each rate is plus a 5.5% agent fee by default.
StepCallPrice
Each drawing clipPOST /v1/videosBy model, at provider list × 1.25; see pricing_skus on GET /v1/videos/models
NarrationPOST /v1/tts-1.0/generate$0.0475 per 1,000 characters
JoinPOST /v1/timeline-1.0/render$0.10 per output minute, reserved in whole minutes
Words on screenPOST /v1/video-captions$0.20 per job, for videos up to 60 seconds

What are the limits?

  • No video model accepts a seed, so a redrawn clip comes out different; keep the takes you approve.
  • Stills for generation must be at public HTTPS URLs; the render takes only this workspace's media.sume.com files, such as the clips Sume returned.
  • The render's default frame is 1080×1920, vertical; set output.width and output.height to match 16:9 clips.
  • For the presenter or B-roll style of explainer instead, see How to make an explainer video with AI.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume