AI cartoon video generator from text: how to set the style

Name the cartoon or anime look in your prompt, or draw a styled still first and animate it as the first frame. Sume's video API has no style field.

5 min readSume
All posts

An AI cartoon video generator from text takes a written description and returns a short animated clip in the look you name, such as a flat 2D cartoon, anime, or claymation. Sume's video models have no style setting, so the look comes from three places: the words of your prompt, a first frame already drawn in that style, or style reference images.

Facts below come from Sume's Video generation, Image API, and Models overview docs, read on 2026-09-28. Anything described as current behavior is read from Sume's API code. To restyle footage you already have, see turning a video into anime.

How do I make a cartoon video from a text prompt?

Write the style and the action into one prompt and send it to POST /v1/videos. The request has no style or genre field: prompt is the text description of the video, and duration, resolution, and aspect_ratio are separate fields. Sume's docs suggest details about motion, camera angles, lighting, and scene composition.

Without code, describe the clip in the Agents tab: the agent picks the models and asks before it spends.

  • Name the look in plain visual words: "flat 2D cartoon, thick outlines, bright colors", "anime, cel shading", or "claymation, clay texture, stop-motion look".
  • Add one action and one camera: who moves, how, and where the camera watches from.
  • Every Sume video model except grok-imagine-video-1.5 can run from the prompt alone. That one needs a first frame.
curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: cartoon-fox-001" \
  -d '{
    "model": "wan-3.0",
    "prompt": "Flat 2D cartoon, thick outlines, bright colors: a small orange fox slides down a snowy hill on its belly and tumbles into a snowbank. Side view, the camera follows.",
    "aspect_ratio": "16:9",
    "resolution": "720p",
    "duration": 8
  }'

Should I draw a cartoon still first?

When a character has to look a certain way, yes. Generate the still from text with POST /v1/images, naming the style in the image prompt, pick the one you like, and send it to POST /v1/videos in frame_images as the first_frame. The clip then opens on your drawing, and the model generates the motion after it.

The Image API returns Sume-hosted, signed result URLs, and signed or private URLs are refused as generation inputs, so host the still you keep at a public HTTPS URL first. For a series of shots with one character, Can AI make a video from a story? covers the still-per-shot workflow, and animating manga panels covers art you already have.

Can reference images set the cartoon style?

Yes, on models that take image references. Send one or more images in input_references with type: "image_url". The docs call them reference images for style guidance, used as visual guidance rather than exact frames, so nothing pins the first frame and the clip can drift from them. Reference-to-video API lists how many each model takes.

Don't mix the two routes in one request: if it also carries frame_images, the frames take precedence and the request is treated as image-to-video. In current code the references are then not sent to the model. The three ways to set the look compare like this:

From Video generation, Video Router, and the catalog behind GET /v1/videos/models, read 2026-09-28.
RouteWhat you sendWhat is pinnedModels
Style in the promptprompt onlyNothing: the look comes from your wordsAll except grok-imagine-video-1.5, which needs a first frame
A styled first frameA still in frame_images as first_frameThe clip's first frameAll
Style referencesImages in input_references as type: "image_url"Nothing: visual guidance rather than exact framesAll except kling-3 and grok-imagine-video-1.5

What are the limits?

  • No field sets the style, and neither a first frame nor references guarantee the look holds in every frame. Watch the clip before you use it.
  • One generation is one short clip: 2–30 seconds depending on the model. A cartoon episode is several clips joined in order.
  • A generated clip won't make a character talk in sync: Sume's docs say video models do not lip-sync to generated speech or to a later voice-over.
  • Frame and reference image URLs must be public HTTPS.
  • Each clip is billed by model, at the rates GET /v1/videos/models lists in pricing_skus.

Sources

Related posts

More in Models

All Models posts

Written by Sume