AssemblyAI Universal-3.6 Pro 32 languages vs Sume STT language_code
AssemblyAI streaming defaults to Universal-3.6 Pro with 32 languages; 3.5 Pro has 19. Sume STT 1.0 takes one optional language_code hint and auto-detects.

AssemblyAI's streaming docs say universal-3-6-pro is the default and recommended model with 32 languages, while universal-3-5-pro keeps the same features but covers 19. Sume STT has no model ladder to choose from: callers use sume/stt-1.0, send one optional language_code hint, or omit it for auto-detect.
AssemblyAI's facts are from its streaming model-selection page; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. Streaming and Sume's batch jobs are different products; this is a language-input reading only.
What does AssemblyAI's streaming page list?
The speech_model connection parameter is optional and defaults to universal-3-6-pro. The page also lists universal-3-5-pro, universal-streaming-english (English) and universal-streaming-multilingual (EN, ES, DE, FR, PT, IT). It describes 3.6 Pro as having native code-switching.
| `speech_model` | Languages | Note on the page |
|---|---|---|
universal-3-6-pro | 32 | Default; native code-switching |
universal-3-5-pro | 19 | Previous flagship, still supported |
universal-streaming-english | English | Cost-effective English |
universal-streaming-multilingual | EN, ES, DE, FR, PT, IT | Per-turn multilingual |
What is the Sume equivalent of picking a model?
There is none for STT. The public id is sume/stt-1.0 and provider models stay internal, so you do not choose between model versions. The only language control is language_code, a BCP-47 or provider hint of 2 to 16 characters, optional, with auto-detect when omitted.
How does the hint map?
Where AssemblyAI has you choose a model by language coverage, Sume has you choose a hint or nothing. Send ko when you know the audio is Korean; leave it out when you do not. Sume publishes no language count, so check a sample in your language.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"language_code": "ko"
}'What else shapes a Sume STT request?
duration_seconds (1 to 600) only improves the usage reservation, and a request holds at most 10 minutes of audio. See the Korean hint walkthrough for a worked example.
Sources
Related posts
More in Models
- Cartesia Sonic 2 and Turbo stop October 20: Sume model ids
Cartesia retires sonic-2 and sonic-turbo on October 20, 2026. Sume's TTS Router lists sonic-3.6, 3.5, 3, latest and preview; here is what to send instead.
- Sonic 3.6 written disfluencies: pacing and Sume's emotion guide
Sonic 3.6 shifts pacing for written hesitations like uh. How to write a natural-sounding transcript, and what Sume's optional emotion guide adds.
- Cartesia STT keyterms for brand names, and Sume STT without keyterms
Cartesia STT adds keyterm prompting for brand names and invented words. Sume STT has no keyterm field, so here is how to fix brand names after transcription.
- Creatify Aurora 2.0 max audio length: 59.7 s, and how to split
Creatify Aurora 2.0 takes up to 59.7 s of audio; Aurora v1 takes 5 minutes. Sume Avatar Video accepts 4-60 s per job, so split longer scripts across jobs.
Written by Sume