AI motion poster generator: animate the art, keep the title

A motion poster moves the artwork while the title stays readable. Animate a text-free version of the art, then put the title back on top of the clip.

5 min readSume
All posts

An AI motion poster generator turns a still poster into a short moving one: the artwork moves (drifting fog, a flickering sign, hair in the wind) while the title and credits stay readable. The way to do that is to animate the art, not the words: give an image-to-video model a version of the poster without its text as the first frame, keep the camera still, then put the title back on top of the finished clip.

Facts come from Sume's Video generation, Video frames, Video captions, and Timeline compose docs, read on 2026-09-28; limits marked as current behavior come from Sume's code. For the still poster itself, see AI poster generator.

How do I make a motion poster with AI?

  • Start from the art with no title, credits, or tagline, cropped to the shape you will post. Sume's docs don't say how a frame of another shape is fitted, so crop first.
  • Host it at a public HTTPS URL and send it to POST /v1/videos as the first_frame in frame_images. Send the same image as the last_frame too if the poster should loop; every image-to-video model except grok-imagine-video-1.5 takes one.
  • Prompt small motion and a still camera: what moves, how slowly, and that the framing stays put.
  • Check the clip: Video frames pulls stills at the times you name, unbilled, so you can look for faces or objects that changed.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: motion-poster-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Static camera. Fog drifts slowly across the night street, the neon sign flickers, her coat moves in the wind. The framing stays the same",
    "frame_images": [
      { "type": "image_url", "image_url": { "url": "https://example.com/poster-art-no-text.png" }, "frame_type": "first_frame" },
      { "type": "image_url", "image_url": { "url": "https://example.com/poster-art-no-text.png" }, "frame_type": "last_frame" }
    ],
    "aspect_ratio": "3:4",
    "resolution": "720p",
    "duration": 6
  }'

Why not animate the poster with the title on it?

Because only the frames you send are your picture. Every frame in between is generated, and Sume's docs make no promise about lettering inside a generated clip, so a title can change shape as the clip plays. Product logo warping in image-to-video covers the same problem with labels. A title added afterwards is text you type or a still you have checked, not lettering the video model generates.

How do I put the title back on?

Two ways. A caption job burns text you type over the art: send the clip to POST /v1/video-captions with one cue, { "text": "…", "start": 0, "end": 6 }, for the whole clip, as in How to add text over a video. Or make the title as a still, such as a title card from a Sume image model whose lettering you have proofread (AI image with text), and put it in the same frame as the clip with POST /v1/timeline-1.0/compose: operation: "stack" splits the frame between the still and the art (still on top by default; ratio sets its share), and "overlay" lays the still over the art as a plate at the top, center, or bottom.

From Video captions and Timeline compose, read 2026-09-28. Each rate is plus a 5.5% agent fee by default; see API pricing.
Caption jobTimeline compose
Where the title comes fromThe words you send in cuesA still image of the title
LayoutText over the art; design.placement sets its heightstack shares the frame; overlay lays a plate over the art
Clip it acceptsA public HTTPS video; in current code 60 seconds or less, with an audio streamThis workspace's media.sume.com files; a clip without sound only warns
OutputThe captioned videoOne MP4 as long as the clip, up to 300 seconds, with the still held throughout
Price$0.20 per job, for videos up to 60 seconds$0.02 per job

What are the limits?

  • No video model in Sume's catalog offers 2:3 or 4:5. The portrait shapes are 3:4 (Seedance 2.x, wan-3.0, minimax-h3, minimax-h3-max) and 9:16.
  • One clip runs 2–30 seconds depending on the model.
  • The caption route needs sound in current code: leave generate_audio out on a model with optional sound, since it defaults to on. Leave style out and Latin text gets slam, which in current code shows it in capitals and lays a light dark tint over the frame.
  • The compose still must already be in this workspace on media.sume.com, such as an earlier Sume output; a title designed in another tool goes on in your own editor after you download the clip. The docs don't say whether a transparent PNG stays transparent, so plan for a solid title plate.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume