Check Seedance or Kling dialogue with a video-inspect transcript

Did your Seedance or Kling clip say your line? Run video-inspect with transcribe true on the Sume-hosted clip and compare the words. $0.01 a minute.

5 min readSume
All posts

To check whether a generated Seedance or Kling clip speaks the line you scripted, send the Sume-hosted clip to POST /v1/video-inspect with transcribe: true, then compare the returned transcript.text with your script. The transcript costs $0.01 per audio minute; the probe and stills are unbilled.

Both model families can produce audio, so a quiet clip, a dropped word or a different language is a real failure mode you can catch cheaply before publishing. The tool reads the clip, never edits it, and is a check rather than a guarantee.

Which clips can I inspect?

Video inspect reads one media.sume.com clip owned by your workspace. Sume mirrors generated artifacts there, so the output of a Seedance or Kling job qualifies. There is no open-internet fetch: a clip hosted elsewhere must be imported first with POST /v1/media-imports, and off-host URLs are rejected at admit.

Requests need an Idempotency-Key. The default mode is sync, which waits up to 30 seconds and returns 200 with the finished inspect, or 202 with a job to poll. See the Video inspect reference.

What does the transcript contain?

With transcribe: true, the resource carries transcript with text, a words[] array, optional sentence segments[] and an audio_url. Add segmentation: {"mode": "sentence"} for caption-line shaped segments. language_code, for example en, is a hint; omit it to auto-detect.

Before paying for a transcript on a clip you suspect is silent, probe it with frames: false and read probe.has_audio. A transcript request on a track-less clip fails with inspect_source_has_no_audio.

What does it cost and refuse?

Costs and failure codes from the docs.

video-inspect transcript facts, read 2026-10-02
ItemValue
Probe and stillsUnbilled
Transcript (Sume STT 1.0)$0.01 per audio minute
Reserve without duration_seconds1 minute; hint max 600 s
Source length cap1800 s
language_code without transcribe400 video_inspect_transcribe_required
Clip with no audioinspect_source_has_no_audio

How do I compare it to my script?

This script transcribes a generated clip and reports which scripted words are missing. It is a crude word check, good for catching a dropped line, not for judging accent or timing.

import os, re, requests

url = os.environ["CLIP_URL"]  # a media.sume.com clip
script = "Our new mug keeps coffee hot all day"
r = requests.post(
    "https://api.sume.com/v1/video-inspect",
    json={"video_url": url, "frames": False, "transcribe": True,
          "language_code": "en", "duration_seconds": 60},
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "inspect-" + url[-24:]},
    timeout=60,
)
r.raise_for_status()
body = r.json()
inspect = body.get("video_inspect", body)
heard = set(re.findall(r"[a-z']+", inspect["transcript"]["text"].lower()))
want = re.findall(r"[a-z']+", script.lower())
print("missing:", [w for w in want if w not in heard])

What are the limits of this check?

Speech-to-text can mishear, so a mismatch is a prompt to listen, not proof of a bad clip. Inspect does not return typed scenes or judge meaning; semantic questions belong to other tools, and only when they are listed. If the 30-second sync window passes, you get a 202 and should poll rather than resubmit.

For multilingual dialogue behavior see AI video dialogue in Korean, Japanese and Spanish, and for the prompt side see the Kling 3.0 prompt guide.

Where does this fit in a publishing workflow?

Run the check after generation and before any edit that would be expensive to redo. A clip that fails the word check can be regenerated with a tighter prompt, or you can keep the visuals and replace the voice using a separate audio step. A clip that passes goes on to trimming, captions or timeline assembly.

Because inspect never re-encodes the source, it is safe to run on the original artifact. Keep the returned audio_url if you want a reviewer to listen to just the track, and use segments[] if the transcript will become caption lines later. The hosted MCP tool video_inspect does the same job for agents, with the same required idempotency_key.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume