Reading a vendor-run TTS leaderboard: Gemini 3.8 Flash TTS on VoiceEQ
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. What that does and does not tell you, plus a blind test to run on your script.

Hume's blog lists Gemini 3.8 Flash TTS at the top of the Hume Real-World VoiceEQ leaderboard. That is a benchmark named for and hosted by Hume, a vendor rather than a neutral lab, so read it as one input and confirm it on your own scripts before you change a pipeline.
The Gemini API changelog separately lists gemini-3.8-flash-tts (studio-grade) and gemini-3.8-flash-lite-tts as generally available as of Sept 22, so both tiers are callable.
What the sources say
Two statements, from two different pages, read on 2026-10-03.
| Claim | Who says it | What it supports |
|---|---|---|
| Ranks first on Real-World VoiceEQ | Hume blog | Strong result on that test's own criteria |
gemini-3.8-flash-tts and flash-lite-tts are GA | Gemini API changelog, Sept 22 | Both tiers can be called now |
| Voice design, replication, 150+ voice library | Gemini API changelog, Sept 22 | Voice options exist; says nothing about quality ranking |
Questions to ask of any vendor-run ranking
These apply to every leaderboard, not only this one.
- Who runs it, and do they sell a competing product? Here the host is Hume, whose name is on the leaderboard.
- What is being scored: emotional fit, pronunciation, naturalness, or speed? A board's name does not tell you it ranks pronunciation of your product names; check what it measures.
- Does the listing say which model tier ranked, and with which voice? A tier in the leaderboard and a tier in your budget may differ.
- Is your content like the test's content? Narration for a 30-second ad is not a customer-support call.
A blind test on your own script
Render the same 3 to 5 lines of your real script on each candidate, then shuffle the files so the listener does not know which is which. This script writes a randomized sheet and a separate answer key from a folder of files named after each model.
import csv, os, random, shutil, sys
src = sys.argv[1] # folder of <model>__<line>.mp3 files
out = sys.argv[2]
os.makedirs(out, exist_ok=True)
files = sorted(f for f in os.listdir(src) if f.endswith((".mp3", ".wav")))
random.shuffle(files)
with open(os.path.join(out, "key.csv"), "w", newline="") as key, \
open(os.path.join(out, "sheet.csv"), "w", newline="") as sheet:
k, s = csv.writer(key), csv.writer(sheet)
k.writerow(["blind_name", "original"])
s.writerow(["blind_name", "score_1_to_5", "notes"])
for i, name in enumerate(files, 1):
blind = f"clip-{i:02d}{os.path.splitext(name)[1]}"
shutil.copy(os.path.join(src, name), os.path.join(out, blind))
k.writerow([blind, name])
s.writerow([blind, "", ""])
print(f"wrote {len(files)} blind clips to {out}")Scoring and next step
Have two or three people score the sheet before opening the key, then join the scores to the key. Keep the criteria narrow: does it read the brand name correctly, does the pacing match the cut, would you ship it. If you generate through an async API, store each take's job id alongside its file so a winning take can be reproduced from its request rather than from memory. On Sume, jobs are async by default and you read the finished artifact from the result URL, which suits batching a test set in one go.
Sources
Related posts
More in Comparisons
- Reference limits: Veo 3.1, Gemini Omni Flash and Sume Video 1.0
Veo 3.1 takes up to 3 reference images, Omni Flash up to 3 clips of 3 s each, and Sume Video 1.0 takes 1 to 9 images. A table to pick by what you need to pin.
- Video API price per second: Veo, Runway, Luma, MiniMax, Seedance
A vendor-authored Higgsfield table lists Veo 3.1 at $0.40 per second down to MiniMax H3 at $0.08. What 10 seconds costs and why to verify each rate.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume