Hume Octave 2 covers 11 languages: is yours one?
Hume Octave 2 is reported to cover 11 languages. Before you pick a voice engine, check your language, then verify what Sume's tts_create recorded for the job.

A Versely roundup says Hume launched Octave 2 on Oct 1, 2025 as a preview, with 11 languages. The roundup lists them as Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian and Spanish, so the first thing to do is check that yours is one of them. On Sume, text-to-speech is the tts_create tool on hosted MCP, and the docs record the language and engine on the job, so you can audit what ran.
What the roundup reports
These numbers are the roundup's, and Sume has not tested the model. Check Hume's own page before you rely on any of them.
| Item | Reported |
|---|---|
| Launch | Oct 1, 2025, as a preview |
| Languages | 11 (Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian, Spanish) |
| Latency | Generation under 200 ms at launch; a later preview nearer 100 ms time to first byte, excluding network |
| Versus Octave 1 | 40% faster, half the price |
| Features | Voice conversion and phoneme editing |
What Sume records about a speech job
Sume lists tts_create among its paid generation tools on hosted MCP (it needs an idempotency_key; under OAuth it needs mcp:write). A finished text_to_speech job records the model_id (the engine), the voice, the language, the output format and the settings it was synthesized with; each is null if the request did not send it. The Sume docs do not give a per-engine language list, so treat language coverage as something to test.
If you route text-to-speech through Sume, read the job after it finishes and confirm the language field is the one you asked for.
A test for any language claim
- Write a 20-second script in your language with numbers, names and one long word.
- Generate it, then read model_id and language back from the job.
- Transcribe the audio and compare it with the script; the video inspect docs give Sume STT 1.0 a public rate of $0.01 per audio minute, to confirm in
GET /v1/catalog. - Have a native speaker listen. Latency figures do not tell you whether pronunciation is right.
Sources
Related posts
More in Models
- HunyuanVideo 1.5 on 14 GB of VRAM: run locally or call an API
The HunyuanVideo 1.5 repo lists 8.3B parameters, 480p to 1080p and a 14 GB VRAM minimum with offloading. A local-run versus API checklist.
- HunyuanVideo 1.5 SSTA and FP8: or just call an API
HunyuanVideo-1.5 has 8.3B parameters, SSTA for a 1.87x speedup at 720p, FP8 GEMM and a 14GB VRAM floor. Decide whether to run it or call a hosted video API.
- Hy Image 3.5 reference limit: 5 or 20, and Sume's caps
Hy Image 3.5's sources disagree on references: 5 in the announcement, 20 in the API guide. Sume's caps differ by model; read input_references from the catalog.
- Music from a thumbnail: image_url on Sume's music router
Generate a music bed that matches a still: pass one public HTTPS image_url with the prompt to Sume's Music Router, and clear it with null when reusing objects.
Written by Sume