AssemblyAI Universal-3.6 Pro 32 languages vs Sume STT language_code

AssemblyAI streaming defaults to Universal-3.6 Pro with 32 languages; 3.5 Pro has 19. Sume STT 1.0 takes one optional language_code hint and auto-detects.

4 min readSume
All posts

AssemblyAI's streaming docs say universal-3-6-pro is the default and recommended model with 32 languages, while universal-3-5-pro keeps the same features but covers 19. Sume STT has no model ladder to choose from: callers use sume/stt-1.0, send one optional language_code hint, or omit it for auto-detect.

AssemblyAI's facts are from its streaming model-selection page; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01. Streaming and Sume's batch jobs are different products; this is a language-input reading only.

What does AssemblyAI's streaming page list?

The speech_model connection parameter is optional and defaults to universal-3-6-pro. The page also lists universal-3-5-pro, universal-streaming-english (English) and universal-streaming-multilingual (EN, ES, DE, FR, PT, IT). It describes 3.6 Pro as having native code-switching.

AssemblyAI streaming model options, vendor docs read 2026-10-01.
`speech_model`LanguagesNote on the page
universal-3-6-pro32Default; native code-switching
universal-3-5-pro19Previous flagship, still supported
universal-streaming-englishEnglishCost-effective English
universal-streaming-multilingualEN, ES, DE, FR, PT, ITPer-turn multilingual

What is the Sume equivalent of picking a model?

There is none for STT. The public id is sume/stt-1.0 and provider models stay internal, so you do not choose between model versions. The only language control is language_code, a BCP-47 or provider hint of 2 to 16 characters, optional, with auto-detect when omitted.

How does the hint map?

Where AssemblyAI has you choose a model by language coverage, Sume has you choose a hint or nothing. Send ko when you know the audio is Korean; leave it out when you do not. Sume publishes no language count, so check a sample in your language.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-demo-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
    "duration_seconds": 120,
    "language_code": "ko"
  }'

What else shapes a Sume STT request?

duration_seconds (1 to 600) only improves the usage reservation, and a request holds at most 10 minutes of audio. See the Korean hint walkthrough for a worked example.

Sources

Related posts

More in Models

All Models posts

Written by Sume