HeyGen photo avatar hand gestures vs Sume Motion Control
HeyGen's motion_prompt still needs an animation reference for Avatar V photo avatars. On Sume, body motion is a separate Motion Control route, not a prompt.
For HeyGen Avatar V photo avatars, a video can now render straight from the photo, but motion_prompt still requires an animation reference, so hand gestures need an eligible reference look. On Sume, a prompt steers the scene of an Avatar Video, while prompted-style body motion from a reference is a separate Motion Control route.
What changed in HeyGen's August 2026 note?
The changelog says that when an Avatar V photo avatar's group has no eligible digital-twin or curated reference look, video creation renders directly from the photo instead of rejecting the request. That works when reference_look_id is omitted and no eligible look can be selected automatically. It then says motion_prompt still requires an animation reference, and to supply an eligible reference look when you need prompted body motion or hand gestures. (Read 2026-09-30.)
| Case | Result per the changelog |
|---|---|
No reference look, no motion_prompt | Renders from the photo |
motion_prompt without a reference | Still requires an animation reference |
| Wants hand gestures | Supply an eligible reference look |
How does Sume steer an avatar video?
Avatar Video takes scene: { "type": "prompt", "prompt": "..." } for scene direction, and a photo scene as another option. That directs the setting, not limb motion. Current execution supports one resolved avatar per final video; see Avatar videos.
Where does body motion come from on Sume?
The OpenAPI file has a face-control alias, createAvatarV1MotionControl, described as a face-control alias of POST /v1/kling/3.0/motion-control with the same body; created jobs store the public model id kling/3.0/motion-control. On the hosted MCP the tool is kling-motion-control_create, listed in the MCP docs. Motion comes from a reference you supply. Read AI avatar hand gestures and dance moves for the setup.
Which route should I use?
If you need a person speaking a script in a chosen scene, use Avatar Video. If you need specific gestures or dance, use Motion Control with a reference clip. They are separate jobs with their own limits, so check each one before combining them.
Sources
Related posts
More in Models
- HeyGen text to video API: 5-15 s, 768p vs Sume durations
heygen-video-1 makes 5-15 second clips at 480p or 768p from a 5,000-character prompt. Sume lists durations and resolutions per model in its catalog.
- HeyGen reference to video: 9 images, 3 videos, 3 audio
HeyGen's heygen-video-1 reference mode takes 9 images, 3 videos and 3 audio files, 12 in total. Here is how Sume's reference limits compare.
- Hy-Image-3.5-Preview API limits vs Sume's image limits
Tencent's Hy-Image-3.5-Preview takes 256 to 8192 px edges, up to 4096x4096 and 20 references. Sume lists no Hy Image model; its own limits are below.
- Ideogram 4.5 Precise Edit API: mask and reference limits
Ideogram 4.5 Precise Edit takes 4 reference images, 3 when a mask is sent. Sume lists Ideogram V3, and mask_url there is only for ChatGPT Image 2.5.
Written by Sume