Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test
Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.

Hume's blog lists Gemini 3.8 Flash TTS at the top of the Hume Real-World VoiceEQ leaderboard. That is useful as a shortlist signal, but the board is run by a company that also sells a voice model, so confirm it with a blind listening test on your own script before you switch.
How to read a vendor-run leaderboard
A vendor-run benchmark is not automatically wrong. It does mean the questions to ask are different from those for an independent test. The blog post gives the ranking; what you need to find out is how the board was built.
- Who wrote the test prompts, and how close are they to your scripts?
- Is the score about emotional expressiveness, naturalness, or accuracy? A board named VoiceEQ suggests an emotion focus, which may not match a plain product narration.
- Were listeners blind to the model, and how many were there?
- What language and voice were used? A top score in one setting may not carry to another.
A blind test you can run in an afternoon
Take five lines from your real scripts: one plain, one with numbers, one with a brand name, one emotional, one long. Generate each with every model you are comparing, keeping the voice type and speed as similar as the models allow. Strip the model names, shuffle the files, and ask three or four colleagues to rank them. Keep the key in a separate file until the votes are in.
This script writes a shuffled listening sheet and a separate answer key.
import csv
import random
MODELS = ["model_a", "model_b", "model_c"]
LINES = ["plain", "numbers", "brand_name", "emotional", "long"]
def main() -> None:
rng = random.Random(42)
with open("sheet.csv", "w", newline="") as sheet, open("key.csv", "w", newline="") as key:
s, k = csv.writer(sheet), csv.writer(key)
s.writerow(["line", "clip"])
k.writerow(["line", "clip", "model"])
n = 0
for line in LINES:
order = MODELS[:]
rng.shuffle(order)
for model in order:
n += 1
s.writerow([line, f"clip_{n:02d}"])
k.writerow([line, f"clip_{n:02d}", model])
if __name__ == "__main__":
main()Record what you generated
A listening test is only repeatable if you know exactly how each clip was made. On Sume, a completed text-to-speech job records model_id, voice ({ "mode": "id", "id": "…" }), language, output_format, and the generation_config and speed it used, with each setting null when the request did not send it. Save those fields next to the clip so that the winner can be recreated later. See Jobs and results.
If your result disagrees with the leaderboard, trust your result for your content. If it agrees, you have a stronger reason to switch than the board alone.
Sources
Related posts
More in Models
- Gemini Omni Flash GA: extension and interpolation vs Sume's inputs
Gemini Omni Flash is generally available with extension and interpolation between images. What Sume's gemini-omni-flash-1.1 catalog entry documents instead.
- Gemini Omni Flash went GA Aug 27: five checks after a preview ends
Omni Flash entered public preview Jun 30 and went GA Aug 27 with extension and 360p-4K. Re-test these five limits before you reuse numbers from preview runs.
- HunyuanVideo 1.5 on 14 GB of VRAM: run locally or call an API
The HunyuanVideo 1.5 repo lists 8.3B parameters, 480p to 1080p and a 14 GB VRAM minimum with offloading. A local-run versus API checklist.
- Luma's 2026 timeline: Ray3.14, Ray3.2, Scenes and Variants
Luma shipped Ray3.14 in January, Ray3.2 in June, Scenes in August and Variants on Oct 1, 2026. What each added, and why to pin model ids.
Written by Sume