Captions when speakers switch languages: Section 508's (in Spanish)

Section 508 says to mark a language change like (in Spanish) and keep the exact words when speech is untranslated. How to burn that with authored Sume cues.

4 min readSume
All posts

When one video holds two spoken languages, Section 508's captioning guidance asks for a descriptor each time the language changes, for example (in Spanish), and then (in English) when the dialogue switches back. If the other-language speech is not translated for hearing viewers, the captions should show the exact wording in that language, with its own grammar and spelling. Sume's speech-to-captions path takes one optional language hint per job, so for a clip that changes language you author the lines as cues.

The rules come from the Captions and Transcripts page on Section508.gov, reviewed August 2025 and read 2026-10-03.

What are the multilingual rules?

The page splits the case in two, depending on whether the other language is translated for hearing viewers.

Section508.gov, Captions and Transcripts, Captioning Different Languages, read 2026-10-03.
SituationWhat the captions must do
Speech is fully translated for hearing viewers, by dubbing or subtitlesInclude the exact same translation
Language changes mid-videoAlways add a descriptor such as (in Spanish), then (in English) when it switches back
Speech is not translated for hearing viewersInclude the exact wording in that language where possible, with correct grammar, spelling and punctuation
No exact transcription is availableSay what you can, such as "arguing in Korean"

Why not let speech-to-text handle it?

Sume's docs describe language as a speech-to-text hint, such as ko or en, and say omitting it detects the language automatically. The hint tells the recognizer what to expect; it does not mark a switch or add a descriptor. A descriptor is editorial copy, so it has to be authored. That is what cues are for: each cue has text, start and end in seconds, and Sume burns exactly that copy at those times without running speech-to-text.

What about Korean or other non-Latin lines?

Mind the style. The docs say slam, punch and tiktok-green draw in Latin display faces without Hangul glyphs, so Korean copy sent to one of them returns 400 (caption_hangul_text_latin_style). For Korean lines use a Hangul identity such as black-outline. A style you name is rendered as named, and language never selects the style or the font, so set the style yourself. Check the video captions docs for the font list, and for scripts beyond Hangul read the post on other scripts Sume documents.

How do I build the cue list?

Keep one list of segments with a language tag and add the descriptor whenever the tag changes. The script below prints the cue list for a short English and Spanish exchange. Each segment is a tuple of language, text, start and end.

import json

segments = [
    ("English", "Welcome to the shop.", 0.5, 2.0),
    ("Spanish", "Bienvenidos a mi tienda.", 2.4, 4.2),
    ("English", "Thank you. Come in.", 4.6, 6.0),
]

cues, current = [], None
for lang, text, start, end in segments:
    if current is not None and lang != current:
        text = f"(in {lang}) {text}"
    current = lang
    cues.append({"text": text, "start": start, "end": end})

print(json.dumps({"cues": cues}, indent=1))

Add video_url to that body and POST it to /v1/video-captions; a job for a clip up to 60 seconds is $0.20. One edge case: this script marks only changes after the first segment, so if a video opens in a language other than the default, label the first cue by hand.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume