Kling motion control silent clip with a Lyria music bed for a Short
Set keep_original_sound to false, generate a track with the Music Router, and join them on a timeline. Cost for a 12-second clip, and the YouTube rule to know.

To avoid carrying the motion video's soundtrack into a Kling motion control clip, send keep_original_sound: false, which returns a silent clip, then add a new track from the Music Router and join the two with a Timeline render. For a 12-second clip the three steps cost about $2.115 on Sume.
Why bother
keep_original_sound defaults to true. If the motion reference is a dance clip filmed with a popular song playing, the output can inherit that audio. YouTube's help page says a Short over one minute with an active Content ID claim of any type is blocked globally, and that nothing changes for Shorts under one minute (read 2026-10-07). A 12-second clip sits under that line, but a clip you later extend or join into a longer Short does not. Starting from silence avoids the question.
Rights are your responsibility either way. A silent clip only removes the original track; it does not clear the motion video itself.
Step 1: the Kling job
POST /v1/kling/3.0/motion-control takes image_url or an avatar (not both), a required motion_video_url, duration_seconds from 1 to 30, an optional prompt that steers appearance only, and character_orientation of video (default) or image. The price is ceil(duration_seconds) times $0.126 times 1.25, which is $0.1575 per second. It is reserved at admission, captured on completion and refunded on failure. The duration_seconds value is the reservation basis; the motion video's length sets the output length. See the silent clip note for the field in isolation.
Step 2: the bed
POST /v1/music-router/generate with model: "sume/music-auto" and a prompt. It rejects duration and duration_seconds, so put the length in the prompt text: "a 12-second instrumental, no vocals". The price is a fixed $0.125 per generation. Check the length you actually received before the join, and set the Timeline's duration_seconds to match it.
Step 3: the join
A Timeline render needs one audio spine and at least one video slot, with all URLs on media.sume.com. Billing is $0.10 per output minute rounded up, so a 12-second render is $0.10.
| Step | Basis | Cost |
|---|---|---|
| Kling motion control | 12 x $0.1575 | $1.89 |
| Music Router generation | fixed | $0.125 |
| Timeline render | 1 started minute | $0.10 |
| Total | $2.115 |
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: kling-bed-001" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/bed.wav",
"duration_seconds": 12
},
"video": [
{
"source_url": "https://media.sume.com/artifacts/artf_demo/kling-silent.mp4",
"start": 0,
"duration": 12
}
]
}'Picking the motion reference
The output motion comes entirely from motion_video_url, so pick a clip with one clear performer, a steady camera and the full body visible if the still shows a full body. character_orientation defaults to video; image is the other value. The prompt is for appearance only, so do not try to direct movement through it. Strict validation also means a stray model or reference_image_urls field is rejected rather than ignored.
Keep duration_seconds honest. It is the basis for the reservation: declare the length of your motion video, rounded up, and the price follows that value.
What to check
Probe the silent clip with video inspect and frames: false; its probe.has_audio should be false. The bed should be quieter than speech if you add any, and soundtrack.duck_db exists for that, but it needs a real audio spine. For a clip over 30 seconds see the 60-second two-job pattern.
Sources
Related posts
More in Use cases
- Landscape listing photos in a vertical reel: cover, contain or blur?
A 3:2 photo in a 9:16 reel shows only 37.5% of its width with fit cover. Pick the fit per photo in Timeline 1.0: eight photos and a bed cost $0.225 on Sume.
- Live stream avatar: queue pre-rendered reaction clips instead
For a stream overlay, render ten short Sume avatar clips ahead of time and trigger them from chat events. Ten 6-second clips cost $14.70 on plus.
- Lock the look with one image, then animate it on three video models
Make one approved image, then use it as first_frame on wan-3.0, kling-3 and minimax-h3-max. Same start, three motions, one prompt, so you can compare.
- A 2.5% word error rate is 25 wrong words per 1,000: plan the proofread
If a 2.5% WER held on your audio, a 1,000-word transcript would still carry about 25 wrong words. How to proofread and burn captions on Sume with script_text.
Written by Sume