Anam Cara-4 anime and 3D avatar images vs Sume's photo input
Anam says Cara-4 animates anime and 3D character images. Sume's avatar API takes a prompt, profile or public HTTPS reference image and promises no style.
Sume's avatar docs do not list anime or 3D characters as supported or unsupported. They describe three input kinds and one URL rule, and nothing about art style. So try your image, and read the job result before you build on it.
Anam facts are from its Cara-4 post; Sume facts from Create new avatar, read 2026-10-01.
What does Anam claim?
Custom avatars no longer need to come from a conventional photo: Cara-4 can animate a wider range of images, including animated 3D characters or anime.
What does Sume accept as input?
On POST /v1/avatar-1.0/generate the input union is a prompt, a props profile, or a photo with image_url. The image_url must be a fetchable public HTTPS image URL. Localhost, private-network, non-HTTPS and non-image responses are rejected before generation. The handle may include a leading @; Sume stores it normalized without @.
| Question | Anam Cara-4 | Sume avatar API |
|---|---|---|
| Anime or 3D image | Stated as supported | Not addressed in the docs |
| Image delivery | Not in the post | Public HTTPS URL |
| Other inputs | Text-prompt avatar on its pricing page | Prompt or profile |
How do I test a cartoon image?
Create the avatar with the photo input, poll the job, and look at the result before spending on a long script. For a cartoon-style talking clip, see AI talking cartoon character from image.
Sources
Related posts
More in Models
- Anam Cara-4 Director Notes vs Sume scene prompts and silence
Anam Director Notes steer a live avatar with a style and an Expressivity control. Sume directs scenes with a prompt or photo, silence beats and a quality tier.
- Anam Cara-4 portrait 768x1152 vs Sume 9:16 avatar at 720p
Anam Cara-4 renders natively at 768x1152 portrait. Sume's avatar video defaults to 9:16 at a fixed 720p, with plans estimated at 4-60 seconds.
- Dictation API: AssemblyAI cleaned text vs Sume STT word timings
AssemblyAI's Dictation API returns send-ready text. Sume STT returns a transcript with words[] timings and no cleanup flag, so you do the filler removal.
- AssemblyAI text to speech: coming soon, and what Sume has today
AssemblyAI's product menu lists a Text-to-Speech API as coming soon. Sume TTS is available now: what a call takes, the 20000-character cap, and word timings.
Written by Sume