OpenAI audio translations go to English only: other languages

OpenAI's /v1/audio/translations endpoint translates into English only. For Korean or Spanish subtitles, transcribe, translate the text, then burn it.

5 min readSume
All posts

OpenAI's audio translation endpoint only translates into English. Its File transcription guide says the endpoint supports translation into English only, at /v1/audio/translations. If you need Korean, Spanish or any other target, the endpoint cannot do it, so you need a text translation step of your own and a way to put the translated text on the video.

Sume has the second half: it burns caption cues onto a video. It has no translation route, so the middle step stays with a translator you choose.

What does the OpenAI guide say about translation?

Read on 2026-10-03, the guide lists three models: gpt-transcribe (the recommended default), gpt-4o-transcribe-diarize for speaker labels, and whisper-1, which the guide ties to timestamps, translations and subtitles. The English-only translation line sits under that last model's section.

The practical result is a gap for every non-English subtitle target. A Spanish speaker's video can become English text, but an English video cannot become Korean text through that endpoint. The same gap applies to Spanish to Japanese.

Where each step lives for non-English subtitles (read 2026-10-03)
StepOpenAI endpointSume
Speech to source textTranscription, 25 MB file limitvideo inspect with transcribe, $0.01 per audio minute
Translate to English/v1/audio/translationsNo route; translate the text yourself
Translate to another languageNot supported by that endpointNo route; translate the text yourself
Burn translated textNot offeredvideo-captions with cues, $0.20 per job up to 60 seconds

How do you turn source text into target-language captions on Sume?

Start with phrase-sized timing. Sume's video inspect can return gapless sentence segments[] when you send transcribe: true and segmentation.mode: "sentence". Each segment has a start and end, so your translator only has to replace the text and keep the times.

Then send the translated lines as cues to video captions. Cues skip speech-to-text entirely, so the job burns exactly your wording at your times. Each cue takes text, start and end in seconds.

curl -X POST https://api.sume.com/v1/video-captions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: es-cues-001" \
  -d '{"video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4", "style": "slam",
       "cues": [{"text": "Hola, esto es Sume", "start": 0.2, "end": 1.6}]}'

Which caption style fits the target language?

Pick the style by the wording you burn, not by the language field. The docs say language is only a speech-to-text hint and never selects the style or the font. Korean text sent to slam, punch or tiktok-green fails with caption_hangul_text_latin_style, so name a Hangul style such as black-outline for Korean.

Spanish, German or French text works on the Latin styles. For other scripts, read the captions guidance before promising a result: Sume documents Latin and Hangul faces, not every script.

Why keep the translation step visible?

A text translation between two jobs is something a person can read. An end-to-end audio translation is not. If a brand name is mistranslated, you can fix that one cue's text and send the cues again.

Budget it simply: one transcription minute rate per source minute, your translator's price, and $0.20 per caption job per language, according to the video captions page. Check GET /v1/catalog for live prices.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume