Gemini TTS reads stage directions aloud: Sume emotion field
Gemini 3.8 TTS speaks its text field verbatim, so style goes in speech_metadata. On Sume, keep the transcript clean and put delivery in generation_config.

Yes, a verbatim TTS model will read your stage directions aloud, so keep delivery out of the spoken text. Google's Gemini 3.8 TTS guide says the text field is strictly a verbatim transcript and puts style in speech_metadata. On Sume the equivalent is generation_config, which holds emotion, speed and volume apart from the transcript.
Gemini facts are from Google's speech generation guide; Sume fields are from the API reference OpenAPI schema.
How does Gemini 3.8 TTS separate text from style?
The guide says to split instructions by scope. Sustained turn-level delivery, such as emotion, style, pacing and volume, goes in speech_metadata.style with values like "whispered urgently". Point-in-time events, such as a pause or a sigh, go inline in angle brackets, for example <short pause>. Everything else in text is spoken.
Where does the same instruction go on Sume?
The TTS request has a separate transcript (1 to 20000 characters, spoken as written) and an optional generation_config with three controls. Putting "(whispering)" in the transcript would be a transcript change, not a delivery setting, so use the field instead.
| Need | Gemini 3.8 TTS | Sume `generation_config` |
|---|---|---|
| Emotion or style for the whole turn | speech_metadata.style | emotion, a string up to 64 characters |
| Speed | Part of style wording | speed, 0.6 to 1.5 |
| Volume | Part of style wording | volume, 0.5 to 2.0 |
| A single pause or sigh | Inline <short pause> tag | Not a documented field; see the linked post |
What does a Sume request look like?
Send the words only in transcript and the delivery in generation_config. The field values below come from the schema; a router model and a voice selector are required on the full request, so add them as your account uses them.
{
"transcript": "Wait, did you hear that?",
"generation_config": {
"emotion": "whispered, urgent",
"speed": 0.9,
"volume": 1.0
}
}What about pauses inside the sentence?
Point-in-time markers differ by provider, and the Sume schema documents no inline tag syntax. Marker handling is covered in text-to-speech: add a pause. For emotion, speed and volume in practice, see text-to-speech API with emotion, speed and volume.
Sources
Related posts
More in Developers
- Gemini API file size limit: 2GB per file, 20GB storage
Gemini's rate-limits page lists a 2GB input file limit and 20GB file storage under Batch API limits. Sume takes public HTTPS URLs instead of uploaded files.
- Gemini Batch API rate limits vs Sume's 100-run bulk queue
Gemini Batch API allows 100 concurrent batch requests and a 2GB input file. Sume bulk runs queue up to 100 Format runs with concurrency 1 to 16.
- Gemini image_size 1k rejected: use 1K, and Sume resolution tiers
Gemini rejects a lowercase image_size such as 1k; use an uppercase K. Sume's images API takes a resolution tier, and only values the model lists.
- Omni 1.1 Flash start and end frame API on Sume
Omni 1.1 Flash can render between two keyframes, including looping clips. On Sume, Omni takes image_url plus end_image_url; frame_images covers other models.
Written by Sume