D-ID audio talks: 15 MB, 5-10 minutes vs Sume's 4-60 s script
D-ID takes up to 40,000 characters of text or 15 MB of audio for a talk. Sume takes an English script of 4 to 60 seconds. Which input fits (read 2026-10-10).
D-ID's Create a talk page lists text scripts up to 40,000 characters (10,000 excluding SSML tags) and audio files up to 15 MB, with audio length limited to 5 minutes for clips and 10 minutes for talks. Sume's Avatar 1.0 takes a text script or multi-scene plan whose estimated duration is 4 to 60 seconds, in English, from an avatar you created once. If you start from a recording, Sume's separate still-plus-audio routes apply instead.
D-ID's figures are from Create a talk and Create a Video, read 2026-10-10. Sume's are from Generate avatar video and Models.
Inputs side by side
The two products are shaped differently, so the limits do not line up.
| Input | D-ID | Sume Avatar 1.0 |
|---|---|---|
| Image | source_url: HTTPS ARGB jpg/png (talk) | Avatar created once from a prompt, traits or a public HTTPS image_url |
| Text script | Up to 40,000 characters (10,000 excluding SSML); clip minimum 3 characters | script whose estimated duration is 4 to 60 seconds |
| Audio script | URL, up to 15 MB; 5 minutes (clips), 10 minutes (talks) | Not a field on talking-video |
| Voices | TTS from Microsoft, ElevenLabs, Amazon, Google or Azure OpenAI | Sume voices; Avatar 1.0 is English-only |
| Result format | mp4 or mov; clips also webm | MP4, 720p |
Where your audio goes on Sume
Sume does not take an audio URL on POST /v1/avatar-1.0/talking-video. For a recorded voice and a face, the models page lists VEED Fabric 1.0 at POST /v1/veed/fabric-1.0, which takes a still plus audio_url and a measured duration_seconds, and MiniMax H3 Max Lip Sync at POST /v1/minimax/h3-max/lip-sync, which takes the same body with audio of 5 to 14.8 seconds. The repo's pricing fixture describes Fabric as up to 300 seconds of audio, Sume-hosted and under 10 MB, at $0.1875 per audio second in 720p.
So a 5-minute recording that D-ID can take as clip audio fits Fabric's 300-second ceiling exactly: 300 seconds at $0.1875 is $56.25. That price is for the Fabric route, not the Avatar 1.0 talking video, whose tiers are $0.184, $0.245 and $0.55 per second (standard, plus, max) before the product-image surcharge.
Choosing by what you have
- You have a script and want a presenter who is reusable across videos: Sume Avatar 1.0, within 4 to 60 seconds. A 60-second standard clip costs 60 x $0.184 = $11.04.
- You have a studio recording of 5 to 10 minutes: D-ID accepts audio of that length; on Sume, split it, or use Fabric up to 300 seconds.
- You have text longer than a minute: D-ID accepts far more characters, while Sume asks you to split the script into several jobs or scenes, each within the window.
- You need languages beyond English: D-ID lists several TTS providers with language-specific voices; Sume Avatar 1.0 will not do it today.
Planning a long script on Sume
Sume estimates duration from the script, so word count is the lever. If an estimate lands above 60 seconds, the request is rejected, so cut it or split it before you submit, and keep the same avatar handle across the pieces so the presenter stays consistent. Each piece is its own job with its own idempotency key. When a script is long, a quick dry run with the avatar-video preview stage shows first frames before any final render is paid for.
Cost check before you split
Splitting a long script is cheap on Sume when you do it deliberately. Two 55-second standard clips are 110 seconds at $0.184, which is $20.24, while the same words as one request would be rejected for exceeding 60 seconds. Cut at a sentence break, keep the same avatar_handle, and give each piece its own Idempotency-Key, so a retry never bills a piece twice.
The D-ID figures here are limits on what a request may contain, not prices, and I did not read D-ID's price list for this post. Compare cost only after you have fetched the current plan page yourself.
Sources
Related posts
More in Comparisons
- ElevenLabs key scopes, quota and IP allowlist vs Sume key controls
ElevenLabs keys can be scoped, capped and IP-locked. Sume keys fix scopes at creation and add per-run spend caps. What each covers and what Sume docs do not.
- ElevenLabs Music 3 s to 5 min vs Sume Music: how to set length
ElevenLabs Music takes 3 seconds to 5 minutes. Sume Music has no duration field and rejects one: set length in the prompt or trim the result afterwards.
- ElevenLabs Music is cleared for nearly all commercial uses: verify
ElevenLabs says Eleven Music is cleared for nearly all commercial uses, with plan differences in separate terms. Check them against your use before publishing.
- ElevenLabs product list vs Sume API routes: what has no equivalent
Sume's OpenAPI has speech, transcription, music and word timings, but no dubbing, voice isolation, sound effects, voice changer or voice cloning route.
Written by Sume