MiniMax Speech 2.8 pitch and emotions vs Sume TTS controls
MiniMax T2A lists speech-2.8 models, nine emotions, pitch -12 to 12 and 10,000 characters. Sume TTS has speed, volume and emotion text, no pitch.

Which controls does MiniMax have that Sume TTS lacks?
Pitch, a fixed emotion list and a wider speed range. The MiniMax text-to-speech reference lists a pitch setting from -12 to 12, nine named emotions, speed from 0.5 to 2 and volume above 0 up to 10. Sume TTS 1.0 documents speed from 0.6 to 1.5, volume from 0.5 to 2.0 and a free-text emotion string, and no pitch control.
If your script needs a higher or lower register than the voice's natural one, change the voice on Sume; there is no pitch dial.
What does the MiniMax reference list?
The endpoint is POST https://api.minimax.io/v1/t2a_v2. The models named are speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, speech-02-turbo, speech-01-hd and speech-01-turbo. Text must be under 10,000 characters, output can be a url or hex, and sample rates run from 8,000 to 44,100 Hz.
Language handling is a language_boost parameter with a long list of languages plus auto. The page does not say which emotions each model supports, so check the model you call before relying on one.
| Control | MiniMax T2A | Sume TTS 1.0 |
|---|---|---|
| Speed | 0.5 to 2 | 0.6 to 1.5 |
| Volume | above 0 up to 10 | 0.5 to 2.0 |
| Pitch | -12 to 12 | No field |
| Emotion | happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper | Free text, up to 64 characters |
| Text limit | Under 10,000 characters | 1 to 20,000 characters |
| Pause marker | <#x#> in text | No documented marker |
How do you map MiniMax settings onto a Sume request?
Translate by intent, not by number. Speed 1.0 means the same on both; MiniMax 1.3 maps to Sume 1.3, and anything outside 0.6 to 1.5 gets clamped to what Sume accepts, so check that the result still reads well. Volume is a multiplier on both, with a narrower ceiling on Sume.
For emotion, write the mood as a short phrase, for example "calm, reassuring". The Sume field is a guide, not a guaranteed mapping, so listen to samples. The emotion, speed and volume post has ranges that worked in practice.
{
"transcript": "Your appointment is confirmed for Friday at ten.",
"avatar_handle": "narrator",
"language": "en",
"generation_config": {
"speed": 0.9,
"volume": 1.2,
"emotion": "calm, reassuring"
}
}Which one should you use?
Choose MiniMax when a project depends on pitch shifts or the named emotions, and accept its own pause markup and rules. Choose Sume when you want one job-based API that also handles music, transcription, timeline audio and avatar video, so the voice file moves straight into the next step by URL.
Sume's TTS Router lists Cartesia Sonic ids only at the moment (see GET /v1/tts-router/models), so there is no MiniMax engine behind Sume today.
Sources
Related posts
More in Comparisons
- MiniMax <#1.5#> pause markup vs Sume TTS: no pause marker
MiniMax inserts pauses with <#x#> markers from 0.01 to 99.99 seconds. Sume TTS documents no pause marker; here is how to split a script and what to use instead.
- MiniMax Video Agent template API: submit, poll, download
MiniMax Video Agent builds a video from a template_id plus your media and text. The endpoints, statuses, and what Sume offers for template-driven video.
- Mirage Tesseract free local engine vs a hosted avatar video API
Mirage Tesseract is an agent video suite with a free local engine. When does a hosted avatar API like Sume Avatar 1.0 fit better? Differences, with dated facts.
- Mirelo SFX video-to-sound vs how Sume adds sound to video
Mirelo generates sound effects synced to existing video. Sume has no video-to-SFX route; here is what it does offer for sound on a clip, and where each fits.
Written by Sume