AI talking cartoon character: make one from a drawing

Make a cartoon character talk: draw or generate one still, add speech audio, and send both to VEED Fabric 1.0 on Sume. Steps, limits, and what is not promised.

5 min readSume
All posts

To make a cartoon character talk, you need two files: one still image of the character and one audio clip of what it says. On Sume you send both to VEED Fabric 1.0, which animates the still to the audio and returns a talking video. Sume's docs do not say how well non-human faces work, so run one short test with your character before you build a series.

The Sume facts are from the Models overview and the Fabric request schema in the Sume API reference. What VEED says about styles is from its Fabric 1.0 page, read 2026-09-29.

What do I need to make a cartoon talk?

Inputs, from the Sume Models docs, the Sume API reference and VEED's page, read 2026-09-29.
InputWhat it must beLimit
image_urlA public HTTPS still of your character (draw it, or generate it with the Images API)Exactly one visual source per request
audio_urlSpeech on Sume's media host, typically a text-to-speech clipMax 10 MB; other hosts are rejected
duration_secondsThe measured length of the audio1 to 300 seconds

Does Fabric work on a clay, anime or cartoon style?

VEED, which makes the model, says Fabric 1.0 creates talking videos in artistic styles including realistic, clay animation and anime, from text prompts or uploaded images, and shows "a clay character explaining concepts" as an example use. That is VEED's claim about its own model. Sume's docs describe the input only as a still image and do not list which styles or non-human faces work, so treat a first clip as a test: check that the mouth reads clearly on your character.

What are the steps?

Create the still first, keeping the character facing the camera with the mouth visible. Make the audio next: Sume's docs say every on-camera speaking shot is a Fabric call with an accepted still plus text-to-speech, because video models do not lip-sync to generated speech or a later voice-over. Then send the still and audio together.

curl -X POST https://api.sume.com/v1/veed/fabric-1.0 \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "image_url": "https://example.com/clay-fox.png",
    "audio_url": "https://media.sume.com/example/fox-line.mp3",
    "duration_seconds": 8,
    "resolution": "720p"
  }'

How long can the character talk, and what does it cost?

One Sume request takes up to 300 seconds of audio; VEED's own page also states talking videos of up to 5 minutes. Fabric is billed per second of audio at $0.1875 per audio second (720p), with the length rounded up to a whole second. For a longer scene, see lip sync AI for long videos.

Should I use Avatar 1.0 for a cartoon?

Not by default. Avatar 1.0 creates a reusable avatar from a text prompt, from structured traits (ethnicity, sex and age), or from a reference photo, and its talking-video route reads a script. The traits are person traits, and the docs do not say what a drawn character does as a reference photo. For a mascot, the still-plus-audio route above is the one whose input you fully control. For a wordless animated clip, see AI cartoon video from text.

Sources

Related posts

More in Sume Avatar 1.0

All Sume Avatar 1.0 posts

Written by Sume