Sonic 3.6 written disfluencies: pacing and Sume's emotion guide
Sonic 3.6 shifts pacing for written hesitations like uh. How to write a natural-sounding transcript, and what Sume's optional emotion guide adds.

Cartesia says Sonic 3.6 reads written disfluencies as natural hesitation: its example is "It's on, uh, Fifth Street", where the uh shifts pacing and intonation. You write the hesitation into the transcript; no tag is needed. On Sume's TTS Router you can send sonic-3.6 as the model and, in the same request, set an optional emotion guide in generation_config.
Sources: the Cartesia 2026 changelog and Sume's API reference, read 2026-10-01.
What does the changelog say about disfluencies?
It lists in-transcript disfluencies among the Sonic 3.6 changes and says the model adapts intonation, pacing and emotiveness to the context of the transcript, with no SSML tags or explicit instructions. It does not say how many hesitations are too many, so judge by ear.
How should I write a hesitant line?
Use it for dialogue that should sound unscripted, not for narration or ads, and keep it rare: one per sentence at most.
| Plain line | With a written hesitation |
|---|---|
| It's on Fifth Street. | It's on, uh, Fifth Street. |
| The total is forty dollars. | The total is, um, forty dollars. |
What does Sume add around the transcript?
The Sume request has generation_config with emotion, an optional emotion guide of 1 to 64 characters, plus speed in [0.6, 1.5] and volume in [0.5, 2.0]. The model field is required on the Router, and sonic-3.6 is in its enum. Text goes in transcript, up to 20000 characters.
The cited schema does not say how the emotion guide interacts with written disfluencies, so treat the pairing as something to test.
How do I test it?
Generate the plain and the hesitant version with the same settings and compare them side by side. If you want pacing control beyond wording, pause markers and speed are the other lever.
Sources
Related posts
More in Models
- Cartesia STT keyterms for brand names, and Sume STT without keyterms
Cartesia STT adds keyterm prompting for brand names and invented words. Sume STT has no keyterm field, so here is how to fix brand names after transcription.
- Creatify Aurora 2.0 max audio length: 59.7 s, and how to split
Creatify Aurora 2.0 takes up to 59.7 s of audio; Aurora v1 takes 5 minutes. Sume Avatar Video accepts 4-60 s per job, so split longer scripts across jobs.
- Creatify Boreal talking clips vs Sume's still-plus-audio route
Creatify says Boreal's gains are smallest on single-person talking clips. Sume makes every speaking shot from an accepted still plus TTS audio via Fabric.
- D-ID V4 expressives sentiment_id vs Sume's emotion string
D-ID picks a delivery with a sentiment_id preset; Sume TTS takes a free-text emotion string plus a 0.6-1.5 speed multiplier. How each one is set.
Written by Sume