OpenAI audio translations go to English only: other languages
OpenAI's /v1/audio/translations endpoint translates into English only. For Korean or Spanish subtitles, transcribe, translate the text, then burn it.

OpenAI's audio translation endpoint only translates into English. Its File transcription guide says the endpoint supports translation into English only, at /v1/audio/translations. If you need Korean, Spanish or any other target, the endpoint cannot do it, so you need a text translation step of your own and a way to put the translated text on the video.
Sume has the second half: it burns caption cues onto a video. It has no translation route, so the middle step stays with a translator you choose.
What does the OpenAI guide say about translation?
Read on 2026-10-03, the guide lists three models: gpt-transcribe (the recommended default), gpt-4o-transcribe-diarize for speaker labels, and whisper-1, which the guide ties to timestamps, translations and subtitles. The English-only translation line sits under that last model's section.
The practical result is a gap for every non-English subtitle target. A Spanish speaker's video can become English text, but an English video cannot become Korean text through that endpoint. The same gap applies to Spanish to Japanese.
| Step | OpenAI endpoint | Sume |
|---|---|---|
| Speech to source text | Transcription, 25 MB file limit | video inspect with transcribe, $0.01 per audio minute |
| Translate to English | /v1/audio/translations | No route; translate the text yourself |
| Translate to another language | Not supported by that endpoint | No route; translate the text yourself |
| Burn translated text | Not offered | video-captions with cues, $0.20 per job up to 60 seconds |
How do you turn source text into target-language captions on Sume?
Start with phrase-sized timing. Sume's video inspect can return gapless sentence segments[] when you send transcribe: true and segmentation.mode: "sentence". Each segment has a start and end, so your translator only has to replace the text and keep the times.
Then send the translated lines as cues to video captions. Cues skip speech-to-text entirely, so the job burns exactly your wording at your times. Each cue takes text, start and end in seconds.
curl -X POST https://api.sume.com/v1/video-captions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: es-cues-001" \
-d '{"video_url": "https://media.sume.com/artifacts/artf_demo/clean.mp4", "style": "slam",
"cues": [{"text": "Hola, esto es Sume", "start": 0.2, "end": 1.6}]}'Which caption style fits the target language?
Pick the style by the wording you burn, not by the language field. The docs say language is only a speech-to-text hint and never selects the style or the font. Korean text sent to slam, punch or tiktok-green fails with caption_hangul_text_latin_style, so name a Hangul style such as black-outline for Korean.
Spanish, German or French text works on the Latin styles. For other scripts, read the captions guidance before promising a result: Sume documents Latin and Hangul faces, not every script.
Why keep the translation step visible?
A text translation between two jobs is something a person can read. An end-to-end audio translation is not. If a brand name is mistranslated, you can fix that one cue's text and send the cues again.
Budget it simply: one transcription minute rate per source minute, your translator's price, and $0.20 per caption job per language, according to the video captions page. Check GET /v1/catalog for live prices.
Sources
Related posts
More in Comparisons
- OpenRouter video expired status vs Sume job statuses
OpenRouter video jobs can end as expired, 'exceeded maximum time to live'. Sume has queued, processing, completed, failed and canceled. Map them in a client.
- OpenRouter video: frame_images and input_references together
Send both frame_images and input_references to an OpenRouter-style video API and Sume treats it as image-to-video. What changes, and how to split the fields.
- OpenRouter video webhook idempotency key vs Sume job_id
OpenRouter's video webhook sends X-OpenRouter-Idempotency-Key as job_id-status. Sume says use job_id alone. How to dedupe both, with a runnable verifier.
- OpenRouter video unsigned_urls need an API key: Sume too
OpenRouter's unsigned_urls require your API key in the Authorization header, and so does Sume's content endpoint. Why a browser video tag fails and a safe fix.
Written by Sume