D-ID API lists V4, V3 and V2 avatars; which to evaluate against Sume
D-ID's overview names V4 Expressive, V3 Pro, V3 Instant and V2 photo avatars. Sume's Avatar 1.0 takes prompt, profile and photo inputs; here is the mapping.
D-ID's API overview lists V4 Expressive Avatars (Full-HD with dynamic expressions and sentiment control), V3 Pro Avatars (Full-HD), V3 Instant Avatars (custom avatars from short videos without training), V2 Avatars (photo-based, from scripts), Video Translate (speech translation with voice cloning and lip-sync) and Agents (realtime). Sume's Avatar 1.0 is one pipeline: create an avatar from a prompt, a profile or a photo, then render talking videos from it. So the question is which D-ID line your use case sits in before you compare anything.
D-ID source: Get started, read 2026-10-04. Sume sources: Create new avatar, Generate avatar video and Face swap (Beta).
A mapping table
This is a rough mapping by job to be done, not a feature-for-feature claim. I did not find resolution or language limits for each D-ID line in the pages read, so none are stated here.
| D-ID line | What the overview says | Nearest Sume route |
|---|---|---|
| V2 Avatars | Photo-based videos using scripts | Avatar from a photo input, then talking-video |
| V3 Instant Avatars | Custom avatars from short videos, no training | No video-based avatar input in the avatar docs; use photo, prompt or profile |
| V3 / V4 avatars | Full-HD, expressions and sentiment | Talking-video at 720p, quality standard, plus or max |
| Video Translate | Translation with voice cloning and lip-sync | Not the same product; Sume has a face swap Beta and captions |
| Agents | Interactive realtime avatars | Not offered; Sume renders clips |
What Sume does not do here
Sume's avatar docs do not describe live, interactive avatars, and its talking video renders at 720p today. If you need realtime conversation or output above 720p, say so in your vendor review and keep that part with the vendor that provides it.
What to test side by side
Run the same 20-second script and the same face on both. Compare the first two seconds of the mouth, the eyes, and the end of the clip, where artefacts accumulate.
- Same script, same aspect ratio, same background idea.
- Judge with sound on and off.
- Time the whole path from request to usable file.
- Check the licence for the likeness you upload.
Bottom line
Treat the D-ID overview as a menu and pick the line first. For script-driven presenter clips from a reusable avatar, Sume's two-step flow is the comparison to run; for realtime or translation, it is not.
Sources
Related posts
More in Comparisons
- Edits First Draft is iOS only: the Android and web route
Instagram's Edits First Draft is iOS only, per a secondary source. On Android or web, cut clips with Sume video trim at $0.02 per job and join them in Timeline.
- Eleven v4 inline tags like [whispers] vs Sume's emotion guide field
Eleven v4 steers delivery with inline tags in the text. Sume TTS has a separate emotion string, plus speed and volume. How to port a tagged script.
- Eleven v4 Turbo 150 ms first speech vs Sume async TTS jobs
ElevenLabs quotes about 150 ms to first speech for v4 Turbo. Sume TTS is an async job, built for finished narration. Which one fits your use?
- ElevenLabs Music for a TV spot: two pages that do not agree
ElevenLabs music page says self-serve plans exclude film, TV and games; the API docs say cleared for film and TV. Get this in writing before a broadcast ad.
Written by Sume