Which languages? Sume's language fields vs MAI's 23 and 60 counts
Microsoft states 23 languages for MAI-Voice-2.1 and 60 for MAI-Transcribe-2-Streaming. Sume publishes no count; here are its language fields.

Microsoft states 23 languages across 26 locales for MAI-Voice-2.1 and 60 languages for MAI-Transcribe-2-Streaming (read 2026-10-08). Sume's docs publish no language count for its speech features, so the honest answer is to test your language on a short sample. What Sume does publish is where language enters each request.
Where a language goes on Sume
On the speech-to-text surfaces, language is a hint rather than a gate.
| Surface | Field | What it does |
|---|---|---|
| Speech to text | language_code (2-16 chars) | Optional hint. Without it, the transcription auto-detects. |
| Video inspect transcript | language_code | Same hint, only valid with transcribe: true; otherwise 400 video_inspect_transcribe_required. |
| Video captions | language | Hint for speech-to-text, for example ko or en. It never selects the caption style or font. |
| Text to speech | language (2-16 chars) | Optional field on the generate body (2 to 16 characters); the body also carries a confirm_language_mismatch flag. |
A sample-first check
Because there is no published list, cut 20 seconds of real audio in the language and run it before committing a batch. Detach a range, then transcribe it. The range option keeps the test to a few cents.
curl -X POST https://api.sume.com/v1/audio-detach \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: lang-sample-001" \
-d '{
"video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
"range": { "start": 0, "end": 20 },
"sample_rate": 16000,
"channels": "mono"
}'Reading the Microsoft counts
The two Microsoft numbers describe different products: 23 languages is voice output, 60 is transcription. A video that is spoken in one language and captioned in another touches both lists, and on Sume touches your own translation step, which is outside the API.
For captions, Video captions documents Korean-specific styles, and rejects Korean text on Latin-only styles with a 400 rather than rendering empty boxes.
A cheaper sample
If the recording is long, cut the sample first with video trim (start 0, duration 20, $0.02), then run video inspect on the trimmed file with frames: false, transcribe: true, duration_seconds: 20 and your language_code. Twenty seconds of audio reserves a third of a cent at $0.01 per audio minute, plus the inspect job's own compute.
Sources
Related posts
More in Developers
- Which Sume API errors to retry and which to stop on: Node wrapper
A retry policy by error.code for Sume submits: retry rate_limited, queue_full, provider_capacity_exceeded; stop on 400, 401, 402, 409. Node wrapper with jitter.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
Written by Sume