Which languages? Sume's language fields vs MAI's 23 and 60 counts

Microsoft states 23 languages for MAI-Voice-2.1 and 60 for MAI-Transcribe-2-Streaming. Sume publishes no count; here are its language fields.

3 min readSume
All posts

Microsoft states 23 languages across 26 locales for MAI-Voice-2.1 and 60 languages for MAI-Transcribe-2-Streaming (read 2026-10-08). Sume's docs publish no language count for its speech features, so the honest answer is to test your language on a short sample. What Sume does publish is where language enters each request.

Where a language goes on Sume

On the speech-to-text surfaces, language is a hint rather than a gate.

From the Sume docs and the API schema
SurfaceFieldWhat it does
Speech to textlanguage_code (2-16 chars)Optional hint. Without it, the transcription auto-detects.
Video inspect transcriptlanguage_codeSame hint, only valid with transcribe: true; otherwise 400 video_inspect_transcribe_required.
Video captionslanguageHint for speech-to-text, for example ko or en. It never selects the caption style or font.
Text to speechlanguage (2-16 chars)Optional field on the generate body (2 to 16 characters); the body also carries a confirm_language_mismatch flag.

A sample-first check

Because there is no published list, cut 20 seconds of real audio in the language and run it before committing a batch. Detach a range, then transcribe it. The range option keeps the test to a few cents.

curl -X POST https://api.sume.com/v1/audio-detach \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: lang-sample-001" \
  -d '{
    "video_url": "https://media.sume.com/artifacts/artf_demo/talk.mp4",
    "range": { "start": 0, "end": 20 },
    "sample_rate": 16000,
    "channels": "mono"
  }'

Reading the Microsoft counts

The two Microsoft numbers describe different products: 23 languages is voice output, 60 is transcription. A video that is spoken in one language and captioned in another touches both lists, and on Sume touches your own translation step, which is outside the API.

For captions, Video captions documents Korean-specific styles, and rejects Korean text on Latin-only styles with a 400 rather than rendering empty boxes.

A cheaper sample

If the recording is long, cut the sample first with video trim (start 0, duration 20, $0.02), then run video inspect on the trimmed file with frames: false, transcribe: true, duration_seconds: 20 and your language_code. Twenty seconds of audio reserves a third of a cent at $0.01 per audio minute, plus the inspect job's own compute.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume