gpt-4o-mini-tts instructions vs Sume TTS speed, volume, emotion

OpenAI steers gpt-4o-mini-tts with free-text instructions. Sume TTS exposes speed 0.6-1.5, volume 0.5-2 and a short emotion string. How the controls compare.

4 min readSume
All posts

OpenAI's text-to-speech guide says gpt-4o-mini-tts takes an instructions parameter that controls accent, emotion, intonation, speed, tone and whispering, in plain language. Sume TTS 1.0 is narrower: generation_config takes a numeric speed from 0.6 to 1.5, a volume multiplier from 0.5 to 2, and a short emotion string of up to 64 characters. If you want a paragraph of acting direction, OpenAI's control is richer; if you want repeatable numeric knobs, Sume's are easier to pin.

The controls side by side

OpenAI's column is what its guide states; Sume's is from the TTS request schema in the Sume API reference.

Controls, read 2026-10-02
ControlOpenAI gpt-4o-mini-ttsSume TTS 1.0
Style directionFree-text instructionsemotion string, 1 to 64 characters
SpeedDescribed inside instructionsspeed 0.6 to 1.5
LoudnessNot a listed parameter on the guidevolume 0.5 to 2
Voices13 built-in voicesVoice from a Sume avatar or a voice id
Output formatsmp3, opus, aac, flac, wav, pcmmp3, wav or raw container; 8 to 48 kHz

A Sume request with the numeric knobs

Sume reads the voice from an avatar (avatar_id or avatar_handle) or voice.id. Set language for any non-English transcript. The default output is mp3 at 44.1 kHz and 128 kbps.

import os, requests

r = requests.post(
    "https://api.sume.com/v1/tts-1.0/generate",
    headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
             "Idempotency-Key": "tts-demo-001"},
    json={"transcript": "Welcome back. Today we cover three updates.",
          "avatar_id": os.environ["SUME_AVATAR_ID"],
          "generation_config": {"speed": 0.95, "emotion": "warm"},
          "output_format": {"container": "mp3", "sample_rate": 44100,
                            "bit_rate": 128000}},
)
print(r.status_code, r.json())

Practical rule

Because a numeric speed and volume do not drift with wording, they suit batches of lines that must sound consistent. For one-off expressive reads, test several short emotion strings and keep the artifact you like rather than expecting identical output on every call.

Worked example

A practical way to compare is to read one paragraph in three moods and see which service gets closest.

  • Pick a 60-word script and three target reads: calm, urgent and warm.
  • On OpenAI, write each as an instructions string and keep the voice fixed.
  • On Sume, keep speed at 1.0 and vary emotion among short strings; then try speeds 0.9 and 1.1.

Checklist before you commit

Pin the numeric values you settle on in your own config so every future run uses them. Sume's generation_config limits are speed 0.6 to 1.5, volume 0.5 to 2, emotion up to 64 characters.

  • Check that your non-English text sets language on Sume.
  • Disclose AI voices to listeners where required; OpenAI's guide asks for this.
  • Save the artifact you like instead of regenerating.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume