One 9-second dance on four brand avatars: Kling motion control, $5.67
Four avatars, one 9-second driving video: four Kling 3.0 Motion Control jobs at $0.1575 a second, 4 x 9 x 0.1575 = $5.67 on Sume. Request shape and queueing.
Animating four brand avatars with the same 9-second dance takes four Kling 3.0 Motion Control jobs and costs 4 x 9 x $0.1575 = $5.67 on Sume (catalog price $0.1575 per output second, read 2026-10-09). Each job takes one visual source and one driving video, so the reference clip is reused by URL rather than merged into one request.
The cost, avatar by avatar
Each job reserves ceil(duration_seconds) times the per-second rate. For a 9.0-second reference that is exactly 9 seconds. If the dance file is 9.04 seconds, it bills as 10 and the total becomes 4 x 10 x 0.1575 = $6.30, so trim the clip to a whole second before you start.
| Avatars | Reference length | Billed seconds | Total |
|---|---|---|---|
| 1 | 9.00 s | 9 | $1.4175 |
| 4 | 9.00 s | 36 | $5.67 |
| 4 | 9.04 s | 40 | $6.30 |
| 10 | 9.00 s | 90 | $14.175 |
How to send it
Each request is POST /v1/kling/3.0/motion-control (or the POST /v1/avatar-1.0/motion-control alias, same body). It needs motion_video_url, a public HTTPS URL for the dance, and duration_seconds from 1 to 30. For the visual source, send either image_url or an avatar_handle / avatar_id, never both. With four ready avatars, use four handles in four requests and one shared motion URL.
Use one Idempotency-Key per avatar, for example dance-avatar-1 through dance-avatar-4, so a client retry cannot reserve the balance twice. By default the driving video's audio stays in the output (keep_original_sound is true). If the dance has music you do not have rights to reuse, set it to false and add your own track afterwards.
Queueing and checks
Four jobs are well inside the limits on any paid plan: Pro runs 4 generation jobs at once and queues up to 20 more (generation-admission docs). On a Free workspace with 1 slot, three of the four wait as queued, which is normal. Poll each job at GET /v1/jobs/:id/status and fetch /result for the video URL.
Before you scale the batch to a full roster, run one avatar and review it. Motion transfer depends on how well the avatar still matches the framing of the dancer, and character_orientation decides whether the video's or the still's framing wins. A quick single test costs $1.4175 and saves you from paying 10 times for the same mistake.
Scaling to a roster
Ten avatars on the same 9-second dance cost 10 x 9 x $0.1575 = $14.175. A Pro workspace runs four at a time and queues the rest, so the last three start after the first wave finishes. Submit all ten with distinct idempotency keys and poll each job; there is no need to space the requests.
Keep a table of handle, job id, status and result URL. When an avatar fails review, rerun that handle only. At $1.4175 per job, one rerun is cheap, but ten reruns of an unreviewed batch is $14.175 again, which is why the single-avatar test comes first.
Sources
Related posts
More in Use cases
- One model for 1:1, 4:5, 9:16, 16:9 and 21:9: 6 Sume rows do it
Of 10 Sume image rows checked, only six list all of 1:1, 4:5, 9:16, 16:9 and 21:9. FLUX.2 pro does it at $0.0375, so five placements cost $0.1875.
- Photo avatar of a real person: when YouTube's AI label applies
A photo avatar scripted to say new words falls under YouTube's 'real person says something they did not' test. How Sume's three avatar inputs map to it.
- Pinterest holiday 2026: carousel stills from one product photo
Pinterest lists Carousel ads, Gift Badge and visual search for holiday 2026. Make four 2:3 stills from one product photo with the Sume Image API.
- Podcast Video to Three Vertical Clips With Captions: $0.66
Cut three 45-second clips from a 30-minute video podcast, burn captions on each and pay $0.66 on Sume: video trim, standalone captions, and the limits to know.
Written by Sume