Video to anime AI: how to turn a clip into anime
To turn a video into anime with AI, use a video-to-video edit that names the style and what to keep, or restyle one frame and animate it.

To turn a video into anime with AI, give a video-to-video (edit) model the clip and a prompt that names the style and says what to keep, for example “Turn this clip into 2D hand-drawn anime. Keep the same people, movements, and camera.” The model rewrites the clip from that instruction. For tighter control of the look, restyle one frame as a still, approve it, and animate it into a new clip instead.
On Sume, the edit is the video-to-video mode of gemini-omni-flash-1.1. Facts come from the Video Router, Image API, and Video generation docs, the Sume API reference, and Sume's API code, read on 2026-09-28.
How do I convert a video to anime with Sume?
Send the clip's public HTTPS URL as video_url to POST /v1/video-router/generate with model: "gemini-omni-flash-1.1", currently the only model with a video-to-video edit, and describe the style in prompt. Edit a video with a prompt covers the other fields, why the edit uses this route, and how to fetch the result.
curl -X POST https://api.sume.com/v1/video-router/generate \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: anime-edit-001" \
-d '{
"model": "gemini-omni-flash-1.1",
"prompt": "Turn this clip into 2D hand-drawn anime with clean line art and flat cel shading. Keep the same people, movements, and camera.",
"video_url": "https://example.com/dance-clip.mp4",
"resolution": "720p",
"mode": "async"
}'What should the style prompt say?
Describe the look in plain visual terms: the drawing style, the line work, the coloring, and the backgrounds. Then say what must stay: the people, their movements, the camera, and the framing. The same pattern covers other styles:
- Anime: “Turn this clip into 2D hand-drawn anime with clean line art and flat cel shading. Keep the same people, movements, and camera.”
- Cartoon: “Make this video look like a bright 2D cartoon with thick outlines and simple shapes. Keep the timing and framing the same.”
- Watercolor: “Restyle this clip as a soft watercolor animation with painted backgrounds. Keep everything else the same.”
- Clay: “Make this clip look like stop-motion clay animation. Keep the same poses and camera.”
Can I control the anime look more tightly?
Yes, by restyling one frame first. The clip you get is a new take, not your original motion:
- Restyle one frame of the video: send it as a still at a public HTTPS URL, in
input_references, to an image model that takes reference images, with the anime look in the prompt. Photo to painting AI walks through this step. - Approve the still, then host it at a public HTTPS URL: the Image API returns signed links, and video requests refuse signed URLs.
- Send it to
POST /v1/videosas thefirst_frame, with a prompt for the motion. The still fixes how the clip opens; the frames after it are generated.
| Edit the video | Restyle a still, then animate | |
|---|---|---|
| Starts from | Your clip in video_url | One frame, restyled by an image model |
| Motion | Your clip, rewritten from the prompt | Generated from the still and a new prompt |
| Model | gemini-omni-flash-1.1 only | An image model that takes references, then a video model |
| Clip length | Not documented | 2–30 seconds, by model |
Will every frame stay in the style?
Sume's docs make no promise that the style holds in every frame or that faces stay recognizable, so watch the whole clip before you use it. They also give no maximum source length and don't say whether the output keeps the source's length, framing, or original sound. The model always generates audio, with no switch to turn it off.
How much does it cost?
An edit is billed per second of output, by resolution, at the provider's list price × 1.25: $0.125 a second at the default 720p, plus a 5.5% agent fee by default. The still-first route adds an image before the clip; AI image generator API cost and How much does an AI video cost? list the rates.
Sources
Related posts
More in Models
- Wan 3.0 vs Seedance 2.5: 30-second clips, inputs and price
Wan 3.0 and Seedance 2.5 both make clips of up to 30 s with audio, frames and references. Wan starts at 2 s and bills per second, Seedance per token.
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
Written by Sume