Kling 4.0 on-screen text vs Sume burned-in captions
Kling 4.0 renders on-screen text inside the video. Sume burns captions onto a finished clip; language is a speech-to-text hint, and cues add authored text.

Kling 4.0's page says it generates on-screen text in 9 languages, plus emoji, that stays readable as the shot moves. That text is drawn by the video model. Sume's route is different: POST /v1/video-captions burns captions onto an existing public video URL, and its language field is a speech-to-text hint, not a translation setting.
Sume's behavior is from Video captions, read 2026-10-01.
Where does the text come from in each case?
With Kling, the model renders the text as part of the picture while it generates. With Sume's captions, the text is added after the video exists, from speech in the clip, from your script_text, or from authored overlay cues. The docs say to prefer this when you already have a finished clip.
| Question | Kling 4.0 (vendor page) | Sume video captions |
|---|---|---|
| When is text added | During generation | After, on a finished clip |
| Languages | 9, per the page | language hints speech-to-text (ko, en, and so on); omit to auto-detect |
| Exact wording | Prompt-driven | script_text aligns to speech; cues or segments burn authored text |
Does the language field translate my captions?
No. The docs describe language as a speech-to-text hint and say it never selects the style or the font. They describe no translation step. If you want captions in another language, supply the translated copy yourself as cues, each with text, start and end.
How do I burn exact text on a silent clip?
Speech-to-captions needs audible speech; a silent clip fails as caption_no_speech. Pass cues or segments to burn authored overlay copy without speech recognition, as in captions for silent clips. script_text, words, cues and segments are mutually exclusive.
What does a caption job cost and need?
The docs say each accepted standalone job reserves and captures $0.20 for videos up to 60 seconds under the current fixed estimate, and tell you to confirm live pricing in GET /v1/catalog. video_url must be a fetchable public HTTPS video. Korean copy needs a Hangul style; Latin styles return 400. See burn captions onto video.
Sources
Related posts
More in Models
- Kling 4.0 stereo audio vs Sume's generate_audio flag
Kling 4.0 moves to two-channel stereo audio. Sume exposes generate_audio and per-model audio capability; check the catalog for what each model reports.
- Kling 4.0 video extension, forward or backward, done by hand on Sume
Kling 4.0 extends a video forward or backward in one workflow. On Sume you chain clips yourself: pull a frame, pass it as first_frame or last_frame.
- Kling Avatar output length follows audio; Sume H3 Max is 5-14.8 s
On fal, Kling Avatar v2 Pro output duration matches the audio file. On Sume the audio sets length too, within route windows from 5 to 14.8 s up to 60 s.
- Kling Avatar v2 Standard vs Pro: fal tiers vs Sume quality
fal prices Kling Avatar v2 Pro at $0.115/s and Standard at $0.0562/s. Sume Avatar Video has its own quality values: standard, plus and max.
Written by Sume