AssemblyAI speech models: unpinned routing and Sume's fixed STT id

AssemblyAI will route unpinned async requests to Universal-3.5 Pro; pin speech_models to stay put. What changes, and why Sume STT has no model to pin.

4 min readSume
All posts

AssemblyAI's August 1, 2026 changelog says that in a future update, async requests that do not pin a model will be routed to Universal-3.5 Pro. To stay on another model you pin it with the plural parameter, for example "speech_models": ["universal-2"]. Sume's speech-to-text route has no such parameter: callers use the public id sume/stt-1.0 and provider models stay internal.

Vendor details are from the AssemblyAI changelog and Sume details from the API reference, both read 2026-10-01.

What does AssemblyAI say will change?

Universal-3.5 Pro is described as the flagship model with the highest accuracy the vendor has released, across 18 languages. The page lists how requests will be handled once the routing update ships.

AssemblyAI async routing after the announced update, from the changelog read 2026-10-01.
Your requestNew behavior
No speech_model or speech_modelsRouted to Universal-3.5 Pro
Singular speech_model parameterRouted to Universal-3.5 Pro
"speech_models": ["universal"]Upgraded to Universal-3.5 Pro
"speech_models": ["universal-3-pro"]Returns an error
u3-pro-rtRedirected to universal-3-5-pro

What action does the vendor ask for?

The page calls pinning universal-3-pro the only breaking change: before the update ships, either pin universal-3-5-pro or omit speech_models to always run on the latest model. It also says the new model may be priced differently than your current configuration, so check your account rates. To stay on a specific model, pin it with the plural parameter.

What is different on Sume STT?

Sume's STT 1.0 description says provider model ids stay internal and callers use Sume-owned job URLs and the public model id sume/stt-1.0. The request takes a public HTTPS audio_url, an optional language_code, and an optional duration_seconds from 1 to 600. The docs state that provider knobs such as diarize and tag_audio_events are fixed server-side.

Billing is by the audio minute: the rate card lists $0.01 per audio minute.

What do I give up by not choosing the model?

You cannot pin a provider model on Sume, so you cannot hold results steady on an older one either; the route is a stable call shape, not a model selector. If your evaluation depends on one named engine, that is a real limit. If it depends on the response shape, the same call keeps working. The same reasoning applies to retirements elsewhere, covered in the Whisper shutdown post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume