ElevenLabs stability 0.5 and similarity 0.75 vs Sume generation_config

ElevenLabs defaults stability to 0.5 and similarity to 0.75. Sume has no such sliders: generation_config takes only volume, speed and an emotion string.

5 min readSume
All posts

ElevenLabs steers a take with voice settings: stability (default 0.5) and similarity_boost (default 0.75) on Text to Speech, and, since its September 28, 2026 changelog, a similarity field (0 to 1, default 0.75) in the settings object of Text to Dialogue. Sume's TTS has no equivalent sliders. Its generation_config accepts three things: volume (0.5 to 2.0), speed (0.6 to 1.5) and a free-text emotion guide of up to 64 characters.

ElevenLabs facts below are from its Text to Speech convert reference and September 28 changelog, read on 2026-10-02.

What do ElevenLabs stability and similarity control?

The convert reference describes stability as running from lower values that "introduce broader emotional range" to higher values that "can result in a monotonous voice with limited emotion", with a default of 0.5. It describes similarity_boost as how closely the AI should adhere to the original voice when replicating it, default 0.75.

Those are two knobs on how a cloned or library voice behaves. The page does not say what each does to latency or cost, so do not infer any.

What does Sume's generation_config give you?

generation_config is optional and has exactly three fields. volume is a multiplier from 0.5 to 2.0. speed is a multiplier from 0.6 to 1.5. emotion is a string from 1 to 64 characters. An older speed enum (slow, normal, fast) still exists at the top level but is marked deprecated; prefer generation_config.speed.

Nothing in the schema controls how closely the output tracks the voice. The voice is whichever one the avatar or voice.id resolves to, and the engine is fixed on TTS 1.0 or chosen from the Router catalog.

How do the controls compare?

ElevenLabs values from the convert reference and changelog; Sume values from the TTS request schema, read 2026-10-02.
ControlElevenLabsSume
Emotional range or monotonestability, default 0.5emotion string only; no numeric scale
Closeness to the source voicesimilarity_boost, default 0.75; similarity 0 to 1 in Text to DialogueNo field
Speaking rateNot in the fields fetched heregeneration_config.speed, 0.6 to 1.5
LoudnessNot in the fields fetched heregeneration_config.volume, 0.5 to 2.0

How do you repeat a take on Sume?

A completed Sume TTS job records what it was made with: model_id, voice ({ "mode": "id", "id": "..." }), language, output_format, generation_config and speed. Each setting is null when the request did not send it. The Jobs and results page says to read these from the job to make the next line sound the same, which is the closest Sume has to saving a preset.

In practice that means a short loop. Submit the line, listen, read the receipt, and copy generation_config into the next request for the following line, changing only speed when a sentence runs long.

Can you tune a Sume take without sliders?

Yes, but with the script and the three fields rather than with a numeric dial. Punctuation shapes pacing, so shorter sentences read tighter and commas add small pauses. speed handles the overall rate, volume handles level, and emotion is a short guide such as a mood word. Keep the text of the line the same while you vary one field at a time, so you know which change you are hearing.

When the output is going into a video, remember the downstream constraint as well: synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded, and a slower speed lengthens the audio. If a long script sits close to the limit, split it by sentence and join the parts with timeline audio concat instead of slowing the voice further.

What should you not expect from Sume?

If your workflow is tuned around ElevenLabs sliders, treat that tuning as ElevenLabs-specific and re-listen on Sume rather than porting the numbers.

  • No numeric stability or similarity dial; emotion is a text hint, not a measured setting.
  • No guarantee that an emotion string changes the audio in a given way; listen to the result.
  • No speed above 1.5 or below 0.6; slow the pacing in the script itself when you need more.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume