Gemini 3.8 Flash TTS voices for a Sume Timeline cut

Gemini 3.8 Flash TTS went GA with voice design and replication. How to bring a voiceover into Sume as the Timeline audio spine, or make it with tts_create.

4 min readSume
All posts

The Gemini API changelog for September 22, 2026 says Gemini 3.8 Flash TTS and Flash-Lite TTS are generally available with voice design and replication, and that the Voice Library now has 150 or more prebuilt and custom voices. On Sume, a voiceover from any source becomes the audio spine of a Timeline 1.0 cut once it is hosted on media.sume.com.

What Google announced

The changelog line covers three points. Those are the only Gemini facts used here, and nothing on this page tests voice quality or price.

Gemini TTS changelog, Sep 22, 2026 (read 2026-10-03)
ItemReported
ModelsGemini 3.8 Flash TTS and Flash-Lite TTS
StatusGenerally available
FeaturesVoice design and replication
Voice Library150+ prebuilt and custom voices

Two ways to get a spine

You can make the voiceover on Sume with the hosted MCP tool tts_create, a paid tool that needs idempotency_key. The docs do not list its model ids, so read its schema with tools_schema. Or you can bring a file made elsewhere. Timeline requires audio already hosted as a media.sume.com artifact or asset, with no open-internet fetch, so import it first with POST /v1/media-imports, or on MCP use the assets_upload_url and assets_complete flow.

Build the cut

Set audio.url to the hosted voiceover and audio.duration_seconds to its length, which can be 1 to 1800 seconds. Place clips in video[] with start and duration on that spine. video[0].start must be 0, and coverage may trail the spine by at most 0.5 seconds. Run POST /v1/timeline-1.0/plan first to see the duration and an estimated cost without creating a job.

curl -X POST https://api.sume.com/v1/timeline-1.0/render \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: vo-cut-001" \
  -d '{
    "audio": {
      "url": "https://media.sume.com/artifacts/artf_demo/voice.wav",
      "duration_seconds": 24
    },
    "video": [
      {"source_url": "https://media.sume.com/artifacts/artf_demo/intro.mp4", "start": 0, "duration": 8},
      {"source_url": "https://media.sume.com/artifacts/artf_demo/body.mp4", "start": 8, "duration": 16}
    ]
  }'

Several takes

If you recorded the narration in sentences, join them first with Timeline audio concat, up to 20 parts, or pass them as audio.parts[] on the render when the join is only needed inside that render. Use wav when the file will be joined again. Voice replication raises its own consent and rights questions, so make sure you have permission to use any voice you clone.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume