Reading a vendor-run TTS leaderboard: Gemini 3.8 Flash TTS on VoiceEQ

Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. What that does and does not tell you, plus a blind test to run on your script.

4 min readSume
All posts

Hume's blog lists Gemini 3.8 Flash TTS at the top of the Hume Real-World VoiceEQ leaderboard. That is a benchmark named for and hosted by Hume, a vendor rather than a neutral lab, so read it as one input and confirm it on your own scripts before you change a pipeline.

The Gemini API changelog separately lists gemini-3.8-flash-tts (studio-grade) and gemini-3.8-flash-lite-tts as generally available as of Sept 22, so both tiers are callable.

What the sources say

Two statements, from two different pages, read on 2026-10-03.

Sources on Gemini 3.8 Flash TTS (read 2026-10-03)
ClaimWho says itWhat it supports
Ranks first on Real-World VoiceEQHume blogStrong result on that test's own criteria
gemini-3.8-flash-tts and flash-lite-tts are GAGemini API changelog, Sept 22Both tiers can be called now
Voice design, replication, 150+ voice libraryGemini API changelog, Sept 22Voice options exist; says nothing about quality ranking

Questions to ask of any vendor-run ranking

These apply to every leaderboard, not only this one.

  • Who runs it, and do they sell a competing product? Here the host is Hume, whose name is on the leaderboard.
  • What is being scored: emotional fit, pronunciation, naturalness, or speed? A board's name does not tell you it ranks pronunciation of your product names; check what it measures.
  • Does the listing say which model tier ranked, and with which voice? A tier in the leaderboard and a tier in your budget may differ.
  • Is your content like the test's content? Narration for a 30-second ad is not a customer-support call.

A blind test on your own script

Render the same 3 to 5 lines of your real script on each candidate, then shuffle the files so the listener does not know which is which. This script writes a randomized sheet and a separate answer key from a folder of files named after each model.

import csv, os, random, shutil, sys

src = sys.argv[1]            # folder of <model>__<line>.mp3 files
out = sys.argv[2]
os.makedirs(out, exist_ok=True)
files = sorted(f for f in os.listdir(src) if f.endswith((".mp3", ".wav")))
random.shuffle(files)

with open(os.path.join(out, "key.csv"), "w", newline="") as key, \
     open(os.path.join(out, "sheet.csv"), "w", newline="") as sheet:
    k, s = csv.writer(key), csv.writer(sheet)
    k.writerow(["blind_name", "original"])
    s.writerow(["blind_name", "score_1_to_5", "notes"])
    for i, name in enumerate(files, 1):
        blind = f"clip-{i:02d}{os.path.splitext(name)[1]}"
        shutil.copy(os.path.join(src, name), os.path.join(out, blind))
        k.writerow([blind, name])
        s.writerow([blind, "", ""])
print(f"wrote {len(files)} blind clips to {out}")

Scoring and next step

Have two or three people score the sheet before opening the key, then join the scores to the key. Keep the criteria narrow: does it read the brand name correctly, does the pacing match the cut, would you ship it. If you generate through an async API, store each take's job id alongside its file so a winning take can be reproduced from its request rather than from memory. On Sume, jobs are async by default and you read the finished artifact from the result URL, which suits batching a test set in one go.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume