Gemini 3.8 Flash TTS tops Hume's VoiceEQ board: run your own test

Hume's blog lists Gemini 3.8 Flash TTS atop its Real-World VoiceEQ board. Why a vendor-run board is only a lead, and how to run a blind A/B on your script.

4 min readSume
All posts

Hume's blog lists Gemini 3.8 Flash TTS at the top of the Hume Real-World VoiceEQ leaderboard. That is useful as a shortlist signal, but the board is run by a company that also sells a voice model, so confirm it with a blind listening test on your own script before you switch.

How to read a vendor-run leaderboard

A vendor-run benchmark is not automatically wrong. It does mean the questions to ask are different from those for an independent test. The blog post gives the ranking; what you need to find out is how the board was built.

  • Who wrote the test prompts, and how close are they to your scripts?
  • Is the score about emotional expressiveness, naturalness, or accuracy? A board named VoiceEQ suggests an emotion focus, which may not match a plain product narration.
  • Were listeners blind to the model, and how many were there?
  • What language and voice were used? A top score in one setting may not carry to another.

A blind test you can run in an afternoon

Take five lines from your real scripts: one plain, one with numbers, one with a brand name, one emotional, one long. Generate each with every model you are comparing, keeping the voice type and speed as similar as the models allow. Strip the model names, shuffle the files, and ask three or four colleagues to rank them. Keep the key in a separate file until the votes are in.

This script writes a shuffled listening sheet and a separate answer key.

import csv
import random

MODELS = ["model_a", "model_b", "model_c"]
LINES = ["plain", "numbers", "brand_name", "emotional", "long"]


def main() -> None:
    rng = random.Random(42)
    with open("sheet.csv", "w", newline="") as sheet, open("key.csv", "w", newline="") as key:
        s, k = csv.writer(sheet), csv.writer(key)
        s.writerow(["line", "clip"])
        k.writerow(["line", "clip", "model"])
        n = 0
        for line in LINES:
            order = MODELS[:]
            rng.shuffle(order)
            for model in order:
                n += 1
                s.writerow([line, f"clip_{n:02d}"])
                k.writerow([line, f"clip_{n:02d}", model])


if __name__ == "__main__":
    main()

Record what you generated

A listening test is only repeatable if you know exactly how each clip was made. On Sume, a completed text-to-speech job records model_id, voice ({ "mode": "id", "id": "…" }), language, output_format, and the generation_config and speed it used, with each setting null when the request did not send it. Save those fields next to the clip so that the winner can be recreated later. See Jobs and results.

If your result disagrees with the leaderboard, trust your result for your content. If it agrees, you have a stronger reason to switch than the board alone.

Sources

Related posts

More in Models

All Models posts

Written by Sume