Gemini 3.8 Flash TTS voices for a Sume Timeline cut
Gemini 3.8 Flash TTS went GA with voice design and replication. How to bring a voiceover into Sume as the Timeline audio spine, or make it with tts_create.

The Gemini API changelog for September 22, 2026 says Gemini 3.8 Flash TTS and Flash-Lite TTS are generally available with voice design and replication, and that the Voice Library now has 150 or more prebuilt and custom voices. On Sume, a voiceover from any source becomes the audio spine of a Timeline 1.0 cut once it is hosted on media.sume.com.
What Google announced
The changelog line covers three points. Those are the only Gemini facts used here, and nothing on this page tests voice quality or price.
| Item | Reported |
|---|---|
| Models | Gemini 3.8 Flash TTS and Flash-Lite TTS |
| Status | Generally available |
| Features | Voice design and replication |
| Voice Library | 150+ prebuilt and custom voices |
Two ways to get a spine
You can make the voiceover on Sume with the hosted MCP tool tts_create, a paid tool that needs idempotency_key. The docs do not list its model ids, so read its schema with tools_schema. Or you can bring a file made elsewhere. Timeline requires audio already hosted as a media.sume.com artifact or asset, with no open-internet fetch, so import it first with POST /v1/media-imports, or on MCP use the assets_upload_url and assets_complete flow.
Build the cut
Set audio.url to the hosted voiceover and audio.duration_seconds to its length, which can be 1 to 1800 seconds. Place clips in video[] with start and duration on that spine. video[0].start must be 0, and coverage may trail the spine by at most 0.5 seconds. Run POST /v1/timeline-1.0/plan first to see the duration and an estimated cost without creating a job.
curl -X POST https://api.sume.com/v1/timeline-1.0/render \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: vo-cut-001" \
-d '{
"audio": {
"url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
"duration_seconds": 24
},
"video": [
{"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
{"source_url": "https://media.sume.com/artifacts/artf_demo/body.mp4", "start": 8, "duration": 16}
]
}'Several takes
If you recorded the narration in sentences, join them first with Timeline audio concat, up to 20 parts, or pass them as audio.parts[] on the render when the join is only needed inside that render. Use wav when the file will be joined again. Voice replication raises its own consent and rights questions, so make sure you have permission to use any voice you clone.
Sources
Related posts
More in Media tools
- Ideogram background remover at $0.01 vs Sume RMBG
fal lists Ideogram Remove Background at $0.01 per image. On Sume, cutouts go through the rmbg_create tool in hosted MCP; check the catalog for its price.
- Indiegogo video 2-4 min: assemble 30 s clips in Timeline
An Indiegogo main video runs 2-4 minutes. Generated clips top out at 30 s on some Sume models, so assemble them with Timeline 1.0 at $0.10 per output minute.
- LinkedIn video thumbnail under 2 MB: pull the frame with Sume
LinkedIn video ad thumbnails are JPG or PNG up to 2 MB. Pull a still from your clip with Sume's video frames endpoint and check its size before upload.
- LTX-2.5 48 fps option: conform frame rate on Sume
LTX-2.5 offers 24/25 fps or 48/50 fps. Sume's video catalog has no fps field, but Timeline output.fps conforms a clip to 24, 25, 30 or 60.
Written by Sume