Public preview with no SLA: MAI voices vs Sume's sonic-preview beta
MAI voices are public preview with no SLA. Sume's sonic-preview is a beta channel that can change; sonic-latest never points to it. Keep both off production.

Both ecosystems have a preview tier you should keep out of production. Microsoft's MAI voices page says the models are in public preview with no SLA and are not recommended for production; its streaming transcription page says the same. On Sume's TTS router, sonic-preview is described as a provider beta channel whose output and availability can change without notice, and sonic-latest is an alias for sonic-3.6 that is never sonic-preview.
What each page says about the tier
The Microsoft voices page, updated 2026-10-01, states public preview, no SLA, and not recommended for production. The realtime transcription page also lists the model as public preview with no SLA. Neither page promises that voices, styles or regions will stay as they are.
Sume's TTS router catalog lists sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview, and GET /v1/tts-router/models returns them. The catalog notes for sonic-preview add one more limit: it is not compatible with pro voice clones, and those requests fail with voice_model_mismatch.
| Question | MAI voices and streaming | Sume sonic-preview |
|---|---|---|
| Status | Public preview | Provider beta channel |
| Guarantee | No SLA | Output and availability can change |
| Production advice | Not recommended | Use a stable id such as sonic-3.6 |
| Alias that avoids it | Not applicable | sonic-latest, never preview |
| Known limit | Instant cloning needs Limited Access Review | Pro voice clones fail with voice_model_mismatch |
How to choose
Use the preview for auditions and tests. For a series that must sound the same next month, pin a stable id such as sonic-3.6; the alias sonic-latest moves when the provider ships a new stable release, which is the right behavior for a general tool and the wrong one for a continuing narrator.
Limits: this compares the labels each vendor gives its tier, not quality. A preview model may well sound better than a stable one, and that is exactly when the temptation to ship it is strongest.
A rule you can encode
Put the decision in code, so nobody ships a preview model by accident. This guard refuses any model id containing preview when the environment is production, and prints the Sume alias that stays on the stable release.
Keep the id in configuration, not in the request code, and log it with each job so a narration series can always be traced to the model that made it.
import os
def pick_tts_model(requested):
env = os.environ.get("APP_ENV", "dev")
if env == "production" and "preview" in requested:
raise SystemExit(f"{requested!r} is a preview channel; use sonic-latest")
return requested
print(pick_tts_model("sonic-latest"))
print(pick_tts_model("sonic-preview"))
Pre-launch checklist for any preview model
If you decide to use a preview anyway, treat it as a dependency that can disappear.
- Render and archive the audio you publish, so a later change to the model cannot alter a finished video.
- Keep a fallback model id in configuration and test the switch once before launch.
- Record the model id with every job, so a drift in sound can be traced.
- Do not promise a customer a voice you cannot re-render on a stable model.
Sources
Related posts
More in Comparisons
- Put a person in a new scene: H3 Max reference video or Recast?
Recast swaps people in a video you have. Reference-to-video makes a new clip from photos. Inputs and 10-second prices: $1.00 vs $3.75 at 768p.
- Luma Ray 3.2 features and the Sume endpoint for each one
Sume does not run Ray 3.2. Match its keyframes, reframe, face tracking and 20 s clips to the Sume surfaces that exist, with docs-verified limits.
- Ray 3.2 takes 16 keyframes; Sume frame_images takes two
Luma lists up to 16 keyframes per Ray 3.2 clip. Sume's frame_images array uses first_frame and last_frame. Which catalog models list them, per the docs.
- Ray 3.2 tracks 8 faces; Sume swaps 1-4 people with H3 Max Recast
Luma Ray 3.2 lists facial tracking for up to 8 faces. Sume has no tracking output, but h3-max-recast swaps 1-4 people in a clip and Kling drives a still.
Written by Sume