MiniMax Speech 2.8 pitch and emotions vs Sume TTS controls

MiniMax T2A lists speech-2.8 models, nine emotions, pitch -12 to 12 and 10,000 characters. Sume TTS has speed, volume and emotion text, no pitch.

5 min readSume
All posts

Which controls does MiniMax have that Sume TTS lacks?

Pitch, a fixed emotion list and a wider speed range. The MiniMax text-to-speech reference lists a pitch setting from -12 to 12, nine named emotions, speed from 0.5 to 2 and volume above 0 up to 10. Sume TTS 1.0 documents speed from 0.6 to 1.5, volume from 0.5 to 2.0 and a free-text emotion string, and no pitch control.

If your script needs a higher or lower register than the voice's natural one, change the voice on Sume; there is no pitch dial.

What does the MiniMax reference list?

The endpoint is POST https://api.minimax.io/v1/t2a_v2. The models named are speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, speech-02-turbo, speech-01-hd and speech-01-turbo. Text must be under 10,000 characters, output can be a url or hex, and sample rates run from 8,000 to 44,100 Hz.

Language handling is a language_boost parameter with a long list of languages plus auto. The page does not say which emotions each model supports, so check the model you call before relying on one.

Voice controls, MiniMax reference and Sume OpenAPI (read 2026-10-02)
ControlMiniMax T2ASume TTS 1.0
Speed0.5 to 20.6 to 1.5
Volumeabove 0 up to 100.5 to 2.0
Pitch-12 to 12No field
Emotionhappy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisperFree text, up to 64 characters
Text limitUnder 10,000 characters1 to 20,000 characters
Pause marker<#x#> in textNo documented marker

How do you map MiniMax settings onto a Sume request?

Translate by intent, not by number. Speed 1.0 means the same on both; MiniMax 1.3 maps to Sume 1.3, and anything outside 0.6 to 1.5 gets clamped to what Sume accepts, so check that the result still reads well. Volume is a multiplier on both, with a narrower ceiling on Sume.

For emotion, write the mood as a short phrase, for example "calm, reassuring". The Sume field is a guide, not a guaranteed mapping, so listen to samples. The emotion, speed and volume post has ranges that worked in practice.

{
  "transcript": "Your appointment is confirmed for Friday at ten.",
  "avatar_handle": "narrator",
  "language": "en",
  "generation_config": {
    "speed": 0.9,
    "volume": 1.2,
    "emotion": "calm, reassuring"
  }
}

Which one should you use?

Choose MiniMax when a project depends on pitch shifts or the named emotions, and accept its own pause markup and rules. Choose Sume when you want one job-based API that also handles music, transcription, timeline audio and avatar video, so the voice file moves straight into the next step by URL.

Sume's TTS Router lists Cartesia Sonic ids only at the moment (see GET /v1/tts-router/models), so there is no MiniMax engine behind Sume today.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume