D-ID V4 expressives sentiment_id vs Sume's emotion string
D-ID picks a delivery with a sentiment_id preset; Sume TTS takes a free-text emotion string plus a 0.6-1.5 speed multiplier. How each one is set.

In D-ID's V4 expressives, you choose the emotional delivery by sending a sentiment_id with an avatar_id to POST /expressives. Sume has no sentiment preset list: in a TTS request you set generation_config.emotion, a free-text string of up to 64 characters, and generation_config.speed, a multiplier in [0.6, 1.5].
D-ID facts are from its V4 quickstart; Sume facts from the OpenAPI schema and Avatar videos, read 2026-10-01.
What is a D-ID sentiment?
The quickstart describes each sentiment as an emotional state, with examples like Professional, Empathetic, Excited, Friendly and Frustrated. Each one comes with its own voice configuration. You list avatars and their sentiments with GET /expressives/avatars, then the video uses the sentiment's configured voice so tone matches the facial expression.
How does Sume express emotion?
On the TTS request, generation_config has optional volume, speed and emotion. The schema describes emotion only as an optional emotion guide for generation, so treat the wording as a hint and listen to the result. For the talking-video path, each scene in a multi-scene plan has its own voice object, and voice.type: "silence" makes a non-speaking beat that needs a duration.
Preset id or free text: what changes?
| Control | D-ID V4 expressives | Sume TTS |
|---|---|---|
| Emotion | sentiment_id from GET /expressives/avatars | generation_config.emotion, free text, up to 64 characters |
| Voice | Configured per sentiment | Chosen separately by avatar or voice id |
| Speed | Not described on the page | generation_config.speed, 0.6 to 1.5 |
| Volume | Not described on the page | volume, 0.5 to 2.0 |
Which one should I use?
A fixed list is easier to validate and keep consistent across a team. A free string gives you more room but no guarantee about what a given word does, so settle on a short vocabulary of your own and test each term. The practical steps are in AI avatar emotion and speaking speed.
What do I send on Sume?
Send the script and voice selector as usual, add generation_config with the emotion and speed you want, and use the same job flow as any other Sume request: poll status_url and read result_url when result_ready is true.
Sources
Related posts
More in Models
- Deepgram nova-3-pharma vs Sume STT: drug-name transcripts
Deepgram added nova-3-pharma for English drug names. Sume STT has one public model, sume/stt-1.0, so check each drug name against word timings.
- Deepgram Nova-3 de-CH, Lithuanian, Portuguese: language_code on Sume
Deepgram improved Nova-3 for Swiss German, Lithuanian and Portuguese on September 24. What to send as language_code on Sume STT for regional variants.
- ElevenLabs character limits by model vs Sume TTS 20,000
ElevenLabs lists 5,000 characters for v3, 10,000 for v4 and 40,000 for Flash v2.5. Sume TTS 1.0 takes up to 20,000 characters in one request.
- ElevenLabs speech to speech API: what Sume offers instead
ElevenLabs lists speech-to-speech voice changer models. Sume has no voice-to-voice endpoint: transcribe with STT, edit the text, then run TTS.
Written by Sume