Gemini TTS reads stage directions aloud: Sume emotion field

Gemini 3.8 TTS speaks its text field verbatim, so style goes in speech_metadata. On Sume, keep the transcript clean and put delivery in generation_config.

4 min readSume
All posts

Yes, a verbatim TTS model will read your stage directions aloud, so keep delivery out of the spoken text. Google's Gemini 3.8 TTS guide says the text field is strictly a verbatim transcript and puts style in speech_metadata. On Sume the equivalent is generation_config, which holds emotion, speed and volume apart from the transcript.

Gemini facts are from Google's speech generation guide; Sume fields are from the API reference OpenAPI schema.

How does Gemini 3.8 TTS separate text from style?

The guide says to split instructions by scope. Sustained turn-level delivery, such as emotion, style, pacing and volume, goes in speech_metadata.style with values like "whispered urgently". Point-in-time events, such as a pause or a sigh, go inline in angle brackets, for example <short pause>. Everything else in text is spoken.

Where does the same instruction go on Sume?

The TTS request has a separate transcript (1 to 20000 characters, spoken as written) and an optional generation_config with three controls. Putting "(whispering)" in the transcript would be a transcript change, not a delivery setting, so use the field instead.

Delivery controls: Gemini versus Sume, read 2026-10-01.
NeedGemini 3.8 TTSSume `generation_config`
Emotion or style for the whole turnspeech_metadata.styleemotion, a string up to 64 characters
SpeedPart of style wordingspeed, 0.6 to 1.5
VolumePart of style wordingvolume, 0.5 to 2.0
A single pause or sighInline <short pause> tagNot a documented field; see the linked post

What does a Sume request look like?

Send the words only in transcript and the delivery in generation_config. The field values below come from the schema; a router model and a voice selector are required on the full request, so add them as your account uses them.

{
  "transcript": "Wait, did you hear that?",
  "generation_config": {
    "emotion": "whispered, urgent",
    "speed": 0.9,
    "volume": 1.0
  }
}

What about pauses inside the sentence?

Point-in-time markers differ by provider, and the Sume schema documents no inline tag syntax. Marker handling is covered in text-to-speech: add a pause. For emotion, speed and volume in practice, see text-to-speech API with emotion, speed and volume.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume