Fish Audio Drama 3 single-word fix vs Sume sentence segments
Fish Audio says Drama 3 preview can fix a single word. Sume has no word-level repair: regenerate a sentence and join takes with Timeline audio concat.

Fish Audio's post says its Drama 3 preview can "fix only a single word", and its API docs call drama-3-preview a preview model whose behavior and availability may change. Sume has no word-level repair. The nearest workflow is to regenerate one sentence as a new take and splice it in with Timeline audio concat, which joins without re-TTS.
Fish Audio facts are from its API docs and a Fish Audio post; Sume facts from the Timeline audio docs and the OpenAPI schema, read 2026-10-01.
What does Fish Audio claim for Drama 3?
The post says you can describe tone, pacing and character in simple language, shift voice mid-sentence, generate a multi-character scene, or fix a single word. The API docs list drama-3-preview among allowed model values and warn that unrecognized values fall back to s2.1-pro. These are vendor claims; this post did not test them.
What can Sume do about one wrong word?
Sume TTS can return gapless sentence segments[], where each segment ends exactly where the next starts. That gives you sentence-sized units to replace. Regenerate only the sentence containing the error, then join old and new pieces.
| Step | Fish Audio Drama 3 (claimed) | Sume |
|---|---|---|
| Unit you redo | One word | One sentence |
| Split the take | Not described | Sentence segments[], gapless |
| Join the pieces | Not described | Timeline audio concat, 1 to 20 ordered parts |
| Seams | Not described | Sample-domain join: no re-TTS, no added silence |
How do I splice a new take in?
Import the audio so it lives on media.sume.com, then call POST /v1/timeline-1.0/audio with operation: "concat" and ordered parts[]. Each part takes a url and optional source_in and duration, so you can trim the old take around the bad sentence. Produced audio is capped at 1800 seconds.
const body = {
operation: "concat",
parts: [
{ url: "https://media.sume.com/artifacts/example/take1.wav", duration: 12.4 },
{ url: "https://media.sume.com/artifacts/example/fixed-sentence.wav" },
{ url: "https://media.sume.com/artifacts/example/take1.wav", source_in: 15.1 },
],
};
console.log(JSON.stringify(body));Will the fixed sentence match the voice and mood?
Not guaranteed. A new take can differ in delivery, and emotion is only an optional guide. Listen at the seams, and keep speed and volume identical across takes. For longer scripts, see the TTS length limit and chapter splitting.
Sources
Related posts
More in Models
- FLUX.2 flex steps and guidance: BFL has them, Sume does not
BFL lists adjustable steps and guidance only for FLUX.2 flex ($0.06/MP). Sume lists flux.2-flex but rejects unlisted parameters with 400.
- FLUX.2 max grounding search: what it is, what Sume lists
Only FLUX.2 [max] does web-grounded generation at BFL. Sume lists flux.2-pro and flux.2-flex, so grounding is not available there.
- Flux TTS expressivity -2 to 2 vs Sume's emotion guide
Deepgram's Flux TTS expressivity runs -2 to 2 (0 nominal). Sume TTS has no such dial: generation_config takes volume, speed and a free-text emotion guide.
- Gemini 3.1 flash image preview shutdown: Sume model ids
Google lists gemini-3.1-flash-image-preview and gemini-3-pro-image-preview for June 25, 2026 shutdown. Sume callers use catalog ids like google/nano-banana-2.
Written by Sume