Final Cut Pro Generate Captions: US English only, Korean fix
Apple's Generate Captions in Final Cut Pro 12.3 is U.S. English only. For Korean speech, Sume video captions takes a language hint and Hangul caption styles.

Final Cut Pro's Generate Captions, added in 12.3, requires a Mac with Apple silicon and is available in U.S. English only, per Apple's release notes. If your speech is Korean, Sume's POST /v1/video-captions takes a language hint and offers the Hangul style korean-ad, reserving $0.20 USD per job under the docs' current estimate.
Apple claims are from its release notes page; Sume claims from Video captions, read 2026-10-01.
What do Apple's notes say about Generate Captions?
Version 12.3 (June 30, 2026) lists Generate Captions to quickly add subtitles and customize their look. The parenthetical reads: requires a Mac with Apple silicon, available in U.S. English only. The notes for 12.4 (September 29, 2026) were also read; the limit above is the one stated for this feature.
How do I burn Korean captions with Sume?
Send the clip as video_url with language: "ko" and, if you want the CapCut-style Hangul look, style: "korean-ad". The docs describe korean-ad as one short phrase at a time, lower-third, with the spoken word shifting to a heavy weight, and say to pair it with language: "ko".
language is only a speech-to-text hint; the docs say it never selects the style or font. Omit it and detection is automatic.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: ko-captions-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/example/clean.mp4",
"style": "korean-ad",
"language": "ko"
}'How do the two approaches differ?
They are different tools: one lives inside a Mac editor, the other is a job API on a video URL.
| Question | Final Cut Pro 12.3 | Sume video captions |
|---|---|---|
| Language | U.S. English only | language hint such as ko or en; omit to auto-detect |
| Hardware | Mac with Apple silicon | None stated for the caller |
| Output | Subtitles in the editor | Captions burned onto the video |
| Price | Not stated in the notes | $0.20 USD up to 60 seconds, fixed estimate |
What if the clip is silent or the wording is known?
A silent clip fails as caption_no_speech; pass cues with text, start and end and Sume burns that copy without speech-to-text. More on styles is in Korean captions via API.
What should I do next?
Start with one short Korean clip, set language: "ko" and korean-ad, and check the burned result before running a batch. Confirm live pricing in GET /v1/catalog, since the docs call $0.20 a current fixed estimate.
Sources
Related posts
More in Use cases
- Firefly Composite API for product photos vs Sume reference edit
Firefly Composite Operations blend a product photo into a generated scene. On Sume, send the photo as an input reference, with an optional mask_url.
- FLUX Virtual Try-On v2: 4 MP inputs and a Sume reference edit
BFL's vto-v2 keeps inputs up to 4 MP as-is. Sume has no try-on endpoint in these docs; a garment swap is a reference edit with public HTTPS images.
- FTC: actors, dramatizations and scripted AI avatar ads
The FTC Q&A says actors in an obviously fictional dramatization are not giving testimonials, yet could still be deceptive. What that means for avatar ads.
- Is #ad enough in a video? What the FTC Q&A says
The FTC says #Ad may work at the start of a text post but may be too easy to miss in a video. Text in a video must stand out. Sume captions burn it in.
Written by Sume