AI talking cartoon character: make one from a drawing
Make a cartoon character talk: draw or generate one still, add speech audio, and send both to VEED Fabric 1.0 on Sume. Steps, limits, and what is not promised.

To make a cartoon character talk, you need two files: one still image of the character and one audio clip of what it says. On Sume you send both to VEED Fabric 1.0, which animates the still to the audio and returns a talking video. Sume's docs do not say how well non-human faces work, so run one short test with your character before you build a series.
The Sume facts are from the Models overview and the Fabric request schema in the Sume API reference. What VEED says about styles is from its Fabric 1.0 page, read 2026-09-29.
What do I need to make a cartoon talk?
| Input | What it must be | Limit |
|---|---|---|
image_url | A public HTTPS still of your character (draw it, or generate it with the Images API) | Exactly one visual source per request |
audio_url | Speech on Sume's media host, typically a text-to-speech clip | Max 10 MB; other hosts are rejected |
duration_seconds | The measured length of the audio | 1 to 300 seconds |
Does Fabric work on a clay, anime or cartoon style?
VEED, which makes the model, says Fabric 1.0 creates talking videos in artistic styles including realistic, clay animation and anime, from text prompts or uploaded images, and shows "a clay character explaining concepts" as an example use. That is VEED's claim about its own model. Sume's docs describe the input only as a still image and do not list which styles or non-human faces work, so treat a first clip as a test: check that the mouth reads clearly on your character.
What are the steps?
Create the still first, keeping the character facing the camera with the mouth visible. Make the audio next: Sume's docs say every on-camera speaking shot is a Fabric call with an accepted still plus text-to-speech, because video models do not lip-sync to generated speech or a later voice-over. Then send the still and audio together.
curl -X POST https://api.sume.com/v1/veed/fabric-1.0 \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image_url": "https://example.com/clay-fox.png",
"audio_url": "https://media.sume.com/example/fox-line.mp3",
"duration_seconds": 8,
"resolution": "720p"
}'How long can the character talk, and what does it cost?
One Sume request takes up to 300 seconds of audio; VEED's own page also states talking videos of up to 5 minutes. Fabric is billed per second of audio at $0.1875 per audio second (720p), with the length rounded up to a whole second. For a longer scene, see lip sync AI for long videos.
Should I use Avatar 1.0 for a cartoon?
Not by default. Avatar 1.0 creates a reusable avatar from a text prompt, from structured traits (ethnicity, sex and age), or from a reference photo, and its talking-video route reads a script. The traits are person traits, and the docs do not say what a drawn character does as a reference photo. For a mascot, the still-plus-audio route above is the one whose input you fully control. For a wordless animated clip, see AI cartoon video from text.
Sources
Related posts
More in Sume Avatar 1.0
- Can HeyGen use my avatar and voice? What its terms say
On Creator, Pro and Business plans you grant HeyGen an irrevocable license to your content, including to train its models. Ask to opt out by email.
- Creatify vs Synthesia: ad avatars or training videos, compared
Creatify sells avatar video ads on credit plans with separate API plans; Synthesia sells presenter videos in 160+ languages with API from Creator.
- Does AI keep your photos and face? What avatar tools store
If an AI tool makes a reusable avatar, it keeps your face on purpose. How long HeyGen, Synthesia and Sume keep it, and how to get it deleted.
- Does HeyGen do captions? What its video API returns
Yes. HeyGen's v3 video API returns an SRT subtitle file when you set caption, and burns captions into the video when you also set a style.
Written by Sume