Kling motion control keep_original_sound: get a silent clip

keep_original_sound on Sume's Kling 3.0 Motion Control defaults to true, so the driving video's audio rides along. Send false for a silent clip.

4 min readSume
All posts

Set keep_original_sound to false on POST /v1/kling/3.0/motion-control and Sume returns a silent clip. Leave it out and the default is true: the driving video's audio track is kept in the output, per Sume's OpenAPI schema for the request.

The choice matters because Motion Control copies motion, not a voice. The still is animated with the driving video's motion, and the output length follows the driving video, so whatever audio that video carries is the only audio the clip can have.

What does keep_original_sound do?

It is a boolean on the Motion Control request. The schema describes it as: keep the driving video's audio track in the output, default true, set false for a silent clip.

Nothing in the schema says Sume generates new sound, mixes audio, or lines a voice up with the mouth. If the driving video has a song, the clip can carry that song. If it has someone talking, the clip can carry that talk, over a still that was animated by the motion.

Source: Sume OpenAPI and docs, read 2026-10-02. Field text from the Kling 3.0 Motion Control request schema.
SettingResultUse it when
omitted or trueThe driving video's audio track is keptThe soundtrack is part of the reference, such as a dance with its music
falseA silent clipYou will add your own music or voice afterwards

When should I send false?

Send false whenever the clip will get new sound later. Sume's models overview says video models do not lip-sync to generated TTS or to a later voice-over, so a talking face is never a motion clip with narration laid underneath. A silent motion clip is for body movement, gestures, and B-roll, with sound added in an edit.

If you need a face that speaks a script, that is a different route: a still plus audio through VEED Fabric 1.0 or MiniMax H3 Max Lip Sync, covered in MiniMax H3 Max lip sync API.

What does the request look like?

Exactly one visual source (image_url or an avatar id or handle), a public HTTPS motion_video_url of at most 30 seconds, and duration_seconds between 1 and 30. This sketch submits with keep_original_sound off and prints the job envelope; poll data.status_url as described in Jobs and results.

import asyncio, os
import httpx

KEY = os.environ.get("SUME_API_KEY", "")
if not KEY:
    raise SystemExit("set SUME_API_KEY")

async def main():
    body = {
        "image_url": "https://example.com/still.png",
        "motion_video_url": "https://example.com/dance.mp4",
        "duration_seconds": 12,
        "keep_original_sound": False,
    }
    async with httpx.AsyncClient(timeout=60) as c:
        r = await c.post(
            "https://api.sume.com/v1/kling/3.0/motion-control",
            headers={"x-api-key": KEY, "Idempotency-Key": "silent-clip-001"},
            json=body,
        )
        r.raise_for_status()
        print(r.json()["data"]["status_url"])

asyncio.run(main())

What should I check on the result?

Open the finished file and confirm it has no audio stream before you ship it. The schema promises a silent clip for false, but it is cheap to verify once per pipeline, and it tells you whether your editor needs a separate soundtrack step.

Reuse the same Idempotency-Key if you retry a submit, so the retry returns the original job instead of a second paid one.

Sources

Related posts

More in Models

All Models posts

Written by Sume