Gemini 3.8 TTS, v4 Turbo, Scribe v2: which can you call via Sume?
None of the three. Sume's catalog (read 2026-10-09) lists sume/tts-1.0, five Sonic router ids, sume/stt-1.0 and three music ids. What to do with their files.

None of them. Sume's public catalog, read on 2026-10-09, lists no Gemini 3.8 Flash TTS or Flash-Lite TTS, no ElevenLabs v4 Turbo and no Scribe v2. For speech it lists sume/tts-1.0 and the TTS Router ids sonic-3.6, sonic-3.5, sonic-3, sonic-latest and sonic-preview, plus sume/stt-1.0 for transcription and sume/music-auto, lyria-3.5 and lyria-3-pro for music.
This post is a catalog check, not a comparison of quality. Vendor facts are from Google's and ElevenLabs' own pricing pages (read 2026-10-09).
What the catalog lists today
The prices below are Sume list prices after margin (provider list x 1.25), as shown in the public catalog. The TTS Router pass-through models bill on the same character book as TTS 1.0. A request is capped at 20,000 characters, and STT estimates cap at 10 minutes.
| Surface | Model ids | Price |
|---|---|---|
| Text to speech | sume/tts-1.0 | $0.0475 per 1,000 characters |
| TTS Router | sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview | $0.0475 per 1,000 characters |
| Speech to text | sume/stt-1.0 | $0.01 per audio minute |
| Music Router | sume/music-auto, lyria-3.5, lyria-3-pro | $0.125 per generation |
| Audio detach | sume/audio-detach-1.0 | $0.01 per job |
What the vendors list
Google's pricing page lists Gemini 3.8 Flash TTS at $9.00 per million audio tokens through December 31, 2026 ($18.00 after) and Flash-Lite TTS at $6.00 ($12.00 after), at 25 tokens per second. ElevenLabs' page lists v4 Turbo at $0.011 per 1K characters with 72% off until Oct 12 (regular $0.04) and Scribe v2 at $0.22 per hour.
If a vendor model is not in Sume's catalog, you cannot ask Sume to run it. On the routers, an id outside the catalog fails with 400 model_not_found.
What you can still do with their files
Audio from any source can be imported into Sume media. Audio detach pulls the sound out of a video at $0.01 a job, and timeline audio joins or splits Sume-hosted files at $0.01 a job. The video captions job takes your own words and skips transcription.
So the honest answer for a team that has picked a vendor voice is: keep generating there, and use Sume for the cutting, joining, rendering and captioning of the result.
How to re-check
The catalog is public and needs no key: fetch https://api.sume.com/v1/catalog and read the models arrays and model_pricing rows. The TTS Router also has GET /v1/tts-router/models for the engines it routes. Prices and ids can change, so any table in a blog post, including this one, is only good for the date it names.
If a model you need is missing, the honest options are to generate with the vendor and import the files, or to use the closest Sume surface and judge the output by ear.
Nothing here ranks the models. A model missing from Sume's catalog is not a statement about its quality, only about what Sume can run for you today.
Sources
Related posts
More in Models
- An OpenRouter-compatible video API: sume/auto or a pinned model
Sume's POST /v1/videos follows OpenRouter's video generation API field for field. Let sume/auto pick the model, or pin a catalog id like seedance-2.5.
- Image generation API with reference images: POST /v1/images
Send a prompt plus public HTTPS reference images to Sume's POST /v1/images. Pin a catalog model or send sume/auto; the catalog lists each model's limits.
- Video 1.0 and Image 1.0 are retiring soon: move to sume/auto
Sume Video 1.0 and Image 1.0 are retiring soon and already run as aliases for the Auto path. New integrations call /v1/videos or /v1/images with sume/auto.
- Music generation API: the Sume Music Router with Lyria 3.5
Sume's Music Router turns a text prompt into a track via POST /v1/music-router/generate. sume/music-auto picks the engine, Lyria 3.5 today.
Written by Sume