iOS 27 generated subtitles: do you still need burned-in captions?

iOS 27 generated subtitles are an on-device fallback for uncaptioned video, English in the US and Canada. What they leave to you, and how to burn captions.

4 min readSume
All posts

Yes, if the wording or the look of your captions matters. Apple's generated subtitles, previewed in May 2026 for release "later this year", are a viewer-side fallback: the device transcribes the audio of a video that arrives without captions. You do not control the words, the timing or the language, and Apple's announcement limits them to English in the U.S. and Canada. A burned-in caption job from Sume puts your wording into the picture before anyone presses play.

This page reads Apple's accessibility announcement and its WWDC26 session on generated subtitles, then lines both up against what Sume's video captions endpoint documents.

What does Apple say generated subtitles do?

Apple describes transcriptions of spoken audio that appear automatically "when captions or subtitles are not already provided", including clips recorded on iPhone, received from friends and family, or streamed online. The text comes from on-device speech recognition, so it is generated privately, and it covers iPhone, iPad, Mac, Apple TV and Apple Vision Pro. Viewers can change how the subtitles look in the playback menu or in Settings.

  • The WWDC26 session names two sources: speech transcription from audio, and language translation from other subtitles.
  • Authored subtitles are preferred and remain unchanged; generated ones fill the gap when the viewer needs a language the video does not carry.
  • Starting in iOS and macOS 27, English subtitles can be generated from English audio; tvOS and visionOS 27 support that too. Multiple languages generated from English subtitles are listed for iOS and macOS.

How does it compare with a burned-in caption job?

The two solve different problems. Apple's feature helps a viewer who has no captions. A burned-in caption is part of the video file, so it shows up the same way in every player, including the ones that never see Apple's subtitle layer.

Apple columns from the Apple Newsroom announcement and the WWDC26 session, read 2026-10-03; Sume column from the Sume video captions docs, read 2026-10-03.
QuestionApple generated subtitlesSume burned-in captions
Who writes the textThe viewer's device, with on-device speech recognitionSpeech-to-text, aligned to your script_text when you pass it, or your own cues
When it appearsWhen the video has no captions or subtitles of its ownAlways, because the text is drawn into the frames
LanguagesEnglish in the U.S. and Canada, per Applelanguage is a speech-to-text hint (ko, en, and others); omit it to auto-detect
LookThe viewer's subtitle styleYour style and design overrides
PriceApple's pages list none$0.20 per accepted job for videos up to 60 seconds

When is the generated fallback enough?

For a family clip or a quick screen recording, Apple's version is a free convenience, and you should leave it alone. It stops being enough when a wrong word costs you something: a product name, a price, a legal line, or a language outside Apple's launch list. Speech recognition can misspell a brand, and a viewer cannot fix that on your behalf.

It also does not replace a deliberate caption design. If your ad needs the hook line large and centered, or a Korean karaoke look, that is a render-time decision.

What do Apple's pages not say?

Neither page says what happens when a video already has burned-in text but no caption track. Apple's trigger is a video without captions or subtitles provided, and burned-in text is pixels rather than a track. Treat the interaction as unknown until you play your own export on an iOS 27 device, and check that a generated line does not stack over yours. If it does, shipping a real caption track next to the burned-in version is the cleaner route; hard versus soft subtitles covers that choice.

How do I burn captions that match my script?

Pass the finished clip as a public HTTPS video_url and the exact spoken script as script_text. Sume keeps the speech-to-text word timings as the timing source and aligns the burned wording to your script. If alignment fails, the job returns script_alignment_mismatch or script_alignment_failed, and omitting script_text burns the recognized wording instead.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: ios27-caption-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/example/clean.mp4",
    "style": "punch",
    "script_text": "Say hello to the Sume developer platform."
  }'

The job is asynchronous: poll GET /v1/jobs/:id/status and reuse the same Idempotency-Key if a network failure makes you retry, as Jobs and results describes. Clips over 60 seconds need splitting first.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume