Hume Octave 2 covers 11 languages: is yours one?

Hume Octave 2 is reported to cover 11 languages. Before you pick a voice engine, check your language, then verify what Sume's tts_create recorded for the job.

3 min readSume
All posts

A Versely roundup says Hume launched Octave 2 on Oct 1, 2025 as a preview, with 11 languages. The roundup lists them as Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian and Spanish, so the first thing to do is check that yours is one of them. On Sume, text-to-speech is the tts_create tool on hosted MCP, and the docs record the language and engine on the job, so you can audit what ran.

What the roundup reports

These numbers are the roundup's, and Sume has not tested the model. Check Hume's own page before you rely on any of them.

Hume Octave 2, as reported by Versely (secondary source, read 2026-10-03)
ItemReported
LaunchOct 1, 2025, as a preview
Languages11 (Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian, Spanish)
LatencyGeneration under 200 ms at launch; a later preview nearer 100 ms time to first byte, excluding network
Versus Octave 140% faster, half the price
FeaturesVoice conversion and phoneme editing

What Sume records about a speech job

Sume lists tts_create among its paid generation tools on hosted MCP (it needs an idempotency_key; under OAuth it needs mcp:write). A finished text_to_speech job records the model_id (the engine), the voice, the language, the output format and the settings it was synthesized with; each is null if the request did not send it. The Sume docs do not give a per-engine language list, so treat language coverage as something to test.

If you route text-to-speech through Sume, read the job after it finishes and confirm the language field is the one you asked for.

A test for any language claim

  • Write a 20-second script in your language with numbers, names and one long word.
  • Generate it, then read model_id and language back from the job.
  • Transcribe the audio and compare it with the script; the video inspect docs give Sume STT 1.0 a public rate of $0.01 per audio minute, to confirm in GET /v1/catalog.
  • Have a native speaker listen. Latency figures do not tell you whether pronunciation is right.

Sources

Related posts

More in Models

All Models posts

Written by Sume