gpt-4o-mini-tts instructions vs Sume TTS speed, volume, emotion
OpenAI steers gpt-4o-mini-tts with free-text instructions. Sume TTS exposes speed 0.6-1.5, volume 0.5-2 and a short emotion string. How the controls compare.

OpenAI's text-to-speech guide says gpt-4o-mini-tts takes an instructions parameter that controls accent, emotion, intonation, speed, tone and whispering, in plain language. Sume TTS 1.0 is narrower: generation_config takes a numeric speed from 0.6 to 1.5, a volume multiplier from 0.5 to 2, and a short emotion string of up to 64 characters. If you want a paragraph of acting direction, OpenAI's control is richer; if you want repeatable numeric knobs, Sume's are easier to pin.
The controls side by side
OpenAI's column is what its guide states; Sume's is from the TTS request schema in the Sume API reference.
| Control | OpenAI gpt-4o-mini-tts | Sume TTS 1.0 |
|---|---|---|
| Style direction | Free-text instructions | emotion string, 1 to 64 characters |
| Speed | Described inside instructions | speed 0.6 to 1.5 |
| Loudness | Not a listed parameter on the guide | volume 0.5 to 2 |
| Voices | 13 built-in voices | Voice from a Sume avatar or a voice id |
| Output formats | mp3, opus, aac, flac, wav, pcm | mp3, wav or raw container; 8 to 48 kHz |
A Sume request with the numeric knobs
Sume reads the voice from an avatar (avatar_id or avatar_handle) or voice.id. Set language for any non-English transcript. The default output is mp3 at 44.1 kHz and 128 kbps.
import os, requests
r = requests.post(
"https://api.sume.com/v1/tts-1.0/generate",
headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Idempotency-Key": "tts-demo-001"},
json={"transcript": "Welcome back. Today we cover three updates.",
"avatar_id": os.environ["SUME_AVATAR_ID"],
"generation_config": {"speed": 0.95, "emotion": "warm"},
"output_format": {"container": "mp3", "sample_rate": 44100,
"bit_rate": 128000}},
)
print(r.status_code, r.json())Practical rule
Because a numeric speed and volume do not drift with wording, they suit batches of lines that must sound consistent. For one-off expressive reads, test several short emotion strings and keep the artifact you like rather than expecting identical output on every call.
Worked example
A practical way to compare is to read one paragraph in three moods and see which service gets closest.
- Pick a 60-word script and three target reads: calm, urgent and warm.
- On OpenAI, write each as an
instructionsstring and keep the voice fixed. - On Sume, keep
speedat 1.0 and varyemotionamong short strings; then try speeds 0.9 and 1.1.
Checklist before you commit
Pin the numeric values you settle on in your own config so every future run uses them. Sume's generation_config limits are speed 0.6 to 1.5, volume 0.5 to 2, emotion up to 64 characters.
- Check that your non-English text sets
languageon Sume. - Disclose AI voices to listeners where required; OpenAI's guide asks for this.
- Save the artifact you like instead of regenerating.
Sources
Related posts
More in Comparisons
- GPT-6 Sol vs Luna vs Astra: context, prices, cutoffs
GPT-6 Astra, Sol and Luna share a 1,050,000-token window but differ 100x in price. A table from OpenAI's own pages, plus which one Sume runs Formats on.
- gpt-image-2.5 partial image streaming vs Sume's 400
OpenAI streams 0 to 3 partial images for gpt-image-2.5. Sume returns 400 streaming_not_supported for stream: true. What to send instead: a job and a poll.
- gpt-image-2.5 rate tiers (5 to 250 IPM) vs Sume plan queue
OpenAI limits gpt-image-2.5 Flare from 5 images per minute at Tier 1 to 250 at Tier 5. Sume limits processing concurrency by plan and queues the rest.
- gpt-realtime-whisper $0.017 per minute vs Sume STT file jobs
OpenAI lists gpt-realtime-whisper at $0.017 a minute for streaming. Sume STT is $0.01 a minute but handles files, not live audio. Pick by workload.
Written by Sume