Check Seedance or Kling dialogue with a video-inspect transcript
Did your Seedance or Kling clip say your line? Run video-inspect with transcribe true on the Sume-hosted clip and compare the words. $0.01 a minute.

To check whether a generated Seedance or Kling clip speaks the line you scripted, send the Sume-hosted clip to POST /v1/video-inspect with transcribe: true, then compare the returned transcript.text with your script. The transcript costs $0.01 per audio minute; the probe and stills are unbilled.
Both model families can produce audio, so a quiet clip, a dropped word or a different language is a real failure mode you can catch cheaply before publishing. The tool reads the clip, never edits it, and is a check rather than a guarantee.
Which clips can I inspect?
Video inspect reads one media.sume.com clip owned by your workspace. Sume mirrors generated artifacts there, so the output of a Seedance or Kling job qualifies. There is no open-internet fetch: a clip hosted elsewhere must be imported first with POST /v1/media-imports, and off-host URLs are rejected at admit.
Requests need an Idempotency-Key. The default mode is sync, which waits up to 30 seconds and returns 200 with the finished inspect, or 202 with a job to poll. See the Video inspect reference.
What does the transcript contain?
With transcribe: true, the resource carries transcript with text, a words[] array, optional sentence segments[] and an audio_url. Add segmentation: {"mode": "sentence"} for caption-line shaped segments. language_code, for example en, is a hint; omit it to auto-detect.
Before paying for a transcript on a clip you suspect is silent, probe it with frames: false and read probe.has_audio. A transcript request on a track-less clip fails with inspect_source_has_no_audio.
What does it cost and refuse?
Costs and failure codes from the docs.
| Item | Value |
|---|---|
| Probe and stills | Unbilled |
| Transcript (Sume STT 1.0) | $0.01 per audio minute |
| Reserve without duration_seconds | 1 minute; hint max 600 s |
| Source length cap | 1800 s |
| language_code without transcribe | 400 video_inspect_transcribe_required |
| Clip with no audio | inspect_source_has_no_audio |
How do I compare it to my script?
This script transcribes a generated clip and reports which scripted words are missing. It is a crude word check, good for catching a dropped line, not for judging accent or timing.
import os, re, requests
url = os.environ["CLIP_URL"] # a media.sume.com clip
script = "Our new mug keeps coffee hot all day"
r = requests.post(
"https://api.sume.com/v1/video-inspect",
json={"video_url": url, "frames": False, "transcribe": True,
"language_code": "en", "duration_seconds": 60},
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "inspect-" + url[-24:]},
timeout=60,
)
r.raise_for_status()
body = r.json()
inspect = body.get("video_inspect", body)
heard = set(re.findall(r"[a-z']+", inspect["transcript"]["text"].lower()))
want = re.findall(r"[a-z']+", script.lower())
print("missing:", [w for w in want if w not in heard])
What are the limits of this check?
Speech-to-text can mishear, so a mismatch is a prompt to listen, not proof of a bad clip. Inspect does not return typed scenes or judge meaning; semantic questions belong to other tools, and only when they are listed. If the 30-second sync window passes, you get a 202 and should poll rather than resubmit.
For multilingual dialogue behavior see AI video dialogue in Korean, Japanese and Spanish, and for the prompt side see the Kling 3.0 prompt guide.
Where does this fit in a publishing workflow?
Run the check after generation and before any edit that would be expensive to redo. A clip that fails the word check can be regenerated with a tighter prompt, or you can keep the visuals and replace the voice using a separate audio step. A clip that passes goes on to trimming, captions or timeline assembly.
Because inspect never re-encodes the source, it is safe to run on the original artifact. Keep the returned audio_url if you want a reviewer to listen to just the track, and use segments[] if the transcript will become caption lines later. The hosted MCP tool video_inspect does the same job for agents, with the same required idempotency_key.
Sources
Related posts
More in Media tools
- Join voiceover takes into one gapless track with Timeline audio
Concatenate up to 20 Sume-hosted voice takes with Timeline audio, with no seam silence and no re-synthesis, and re-base video starts from the returned offsets.
- Crop 16:9 to 4:5 with video-filter: width 0.45 and an x offset
To crop 16:9 to 4:5 in Sume video-filter, use width 0.45, height 1 and an x between 0 and 0.55. Here is the math, the x choice and the output scale.
- Crop a 16:9 video to a centered 9:16 strip: the crop fractions
For a centered 9:16 crop of a 16:9 video, send crop x 0.3418, y 0, width 0.3164, height 1 to Sume video-filter. 1:1 and 4:5 values are in the table.
- Cut out a subject from a video frame: video frames, then RMBG
Pull a lossless PNG frame with Sume video frames (unbilled), then run RMBG 1.0 on it for a transparent PNG. Sume does not remove backgrounds from a whole video.
Written by Sume