Cloudflare Stream captions API: 12 languages, or upload WebVTT

Cloudflare Stream generates captions for 12 languages. For others, PUT a WebVTT built from a Sume transcript. Rules, status values and a script.

5 min readSume
All posts

Cloudflare Stream can generate captions itself for 12 languages, and for anything else you upload a WebVTT file with a PUT to the video's captions endpoint (read 2026-10-03). Sume fits the second path: transcribe the video with Sume's transcript call, write the segments as WebVTT, and upload that file under the right language tag.

This post lists the languages Cloudflare generates, the rules that make an upload replace a generated track, and the script that sends the file.

Which languages does Cloudflare generate captions for?

Cloudflare's captions page lists Czech, Dutch, English, French, German, Italian, Japanese, Korean, Polish, Portuguese, Russian and Spanish for generated captions (read 2026-10-03). The video must already be uploaded and in a ready state, and you should generate for the language actually spoken in the audio. A video can carry captions in several languages, but each language must be unique: you cannot have two English tracks.

Cloudflare Stream captions behavior, read 2026-10-03
TopicWhat the Cloudflare page saysPractical result
GeneratePOST to .../captions/<LANGUAGE_TAG>/generate; status inprogress, ready or errorWorks only for the 12 listed languages
UploadPUT a file to .../captions/<LANGUAGE_TAG>Use this for every other language
Language tagBCP 47, for example en or en-GBCloudflare builds the player label from the tag
One per languageEach language must be uniqueA PUT to a language you already have replaces that track
Editing a generated trackgenerated becomes false and the auto-generated label suffix goesAn uploaded file is never marked generated
FetchGET .../captions/<LANGUAGE_TAG>/vtt returns the WebVTTHandy for diffing against your Sume output

How do I get a WebVTT file out of Sume?

Use POST /v1/video-inspect with transcribe: true and segmentation.mode: "sentence". It reads one clip that already lives on media.sume.com, so import the video first with POST /v1/media-imports; the public rate for the transcript is $0.01 per audio minute (video inspect). The segments come back gapless, in seconds, and the full script for turning them into a clean VTT file is in the Vimeo post. Cloudflare's page does not state a cue-timing rule like Vimeo's, but a one-millisecond trim between cues is harmless.

Set language_code on the request when you know the language; leaving it out auto-detects. Read the first minute of output before you upload: Sume returns speech-to-text wording, and Cloudflare will display exactly what you send.

Cloudflare lists status values for generated tracks only (inprogress, ready, error). An uploaded file comes back as ready with generated set to false, so your pipeline can treat the PUT response as the confirmation instead of polling. If you want proof the player manifest has the track, fetch the captions list for the video and look for your language entry with status ready.

Keep the Sume job and the Cloudflare video paired in your own records, for example by storing the Sume job id next to the Cloudflare UID, so a later correction can regenerate and re-upload exactly the right track without guessing which transcript went where.

How do I upload it?

The script below sends talk.vtt as a multipart file to the language endpoint and prints the result, which Cloudflare returns with generated: false and a status. It needs an API token with Stream edit rights, your account id and the video UID.

import os, requests

acct, uid, lang = os.environ["CF_ACCOUNT_ID"], os.environ["CF_VIDEO_UID"], "ar"
url = f"https://api.cloudflare.com/client/v4/accounts/{acct}/stream/{uid}/captions/{lang}"
headers = {"Authorization": f"Bearer {os.environ['CF_API_TOKEN']}"}

with open("talk.vtt", "rb") as f:
    r = requests.put(url, headers=headers, files={"file": ("talk.vtt", f, "text/vtt")})
r.raise_for_status()
cap = r.json()["result"]
print(cap["language"], cap["label"], "generated:", cap["generated"], cap["status"])

What should I watch for?

A few details decide whether the upload shows up in the player the way you expect.

  • Cloudflare describes the PUT as creating or replacing a caption file, so re-running the script after a fix overwrites the previous upload.
  • If a generated caption lands in an error state, Cloudflare's advice is to delete it and generate again.
  • Pick the tag deliberately: Cloudflare builds the player label from it, and a regional tag such as en-GB renders a label for British English.
  • The list call also returns generated tracks that are inprogress or in error, so filter on status before you assume a track is live.
  • A sidecar track is switchable by the viewer. If the clip must carry words everywhere it is reposted, burn captions with video captions as well.

Is generating in Cloudflare or uploading from Sume the better route?

For one of the 12 languages and a clean recording, Cloudflare's generation is a single call and needs no second vendor. Choose the Sume route when the language is outside that list, when you want to correct wording in code before publishing, or when the same transcript also feeds a burned-in version for social. Sume STT handles up to 600 seconds per call, so longer videos are transcribed in parts, as the Spotify walk-through shows. If your Cloudflare need is only a short excerpt, compare the clip API with Sume video trim.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume