D-ID API lists V4, V3 and V2 avatars; which to evaluate against Sume

D-ID's overview names V4 Expressive, V3 Pro, V3 Instant and V2 photo avatars. Sume's Avatar 1.0 takes prompt, profile and photo inputs; here is the mapping.

5 min readSume
All posts

D-ID's API overview lists V4 Expressive Avatars (Full-HD with dynamic expressions and sentiment control), V3 Pro Avatars (Full-HD), V3 Instant Avatars (custom avatars from short videos without training), V2 Avatars (photo-based, from scripts), Video Translate (speech translation with voice cloning and lip-sync) and Agents (realtime). Sume's Avatar 1.0 is one pipeline: create an avatar from a prompt, a profile or a photo, then render talking videos from it. So the question is which D-ID line your use case sits in before you compare anything.

D-ID source: Get started, read 2026-10-04. Sume sources: Create new avatar, Generate avatar video and Face swap (Beta).

A mapping table

This is a rough mapping by job to be done, not a feature-for-feature claim. I did not find resolution or language limits for each D-ID line in the pages read, so none are stated here.

D-ID lines and the nearest Sume route (read 2026-10-04)
D-ID lineWhat the overview saysNearest Sume route
V2 AvatarsPhoto-based videos using scriptsAvatar from a photo input, then talking-video
V3 Instant AvatarsCustom avatars from short videos, no trainingNo video-based avatar input in the avatar docs; use photo, prompt or profile
V3 / V4 avatarsFull-HD, expressions and sentimentTalking-video at 720p, quality standard, plus or max
Video TranslateTranslation with voice cloning and lip-syncNot the same product; Sume has a face swap Beta and captions
AgentsInteractive realtime avatarsNot offered; Sume renders clips

What Sume does not do here

Sume's avatar docs do not describe live, interactive avatars, and its talking video renders at 720p today. If you need realtime conversation or output above 720p, say so in your vendor review and keep that part with the vendor that provides it.

What to test side by side

Run the same 20-second script and the same face on both. Compare the first two seconds of the mouth, the eyes, and the end of the clip, where artefacts accumulate.

  • Same script, same aspect ratio, same background idea.
  • Judge with sound on and off.
  • Time the whole path from request to usable file.
  • Check the licence for the likeness you upload.

Bottom line

Treat the D-ID overview as a menu and pick the line first. For script-driven presenter clips from a reusable avatar, Sume's two-step flow is the comparison to run; for realtime or translation, it is not.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume