Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken
Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.

If a TTS transcript contains caption markup, Sume removes it before the voice sees it: a || cue break and an @{...} marker say nothing, and <display|spoken> keeps only the spoken side. The strip runs before validation, billing and storage, so the voice never reads "vertical bar". If nothing speakable is left, for example a transcript of || @{reveal!} ||, the API returns 400 transcript_no_speech with details.field set to transcript.
What gets removed
The rule lives in the API submit path for POST /v1/tts-1.0/generate. A cue break always separates words, so a||b is spoken as two words, and line breaks inside the transcript are kept, because a voice reads them as pauses. Text without any markup comes back byte for byte unchanged.
Captions are the other half of the story: the same markup is meant for caption cue timing, so one script can drive both the voice and the captions without a hand-cleaned copy for each.
| You send | Voice speaks | Note |
|---|---|---|
| Hello||world | Hello world | Cue break becomes a space |
| Big sale @{pop!} today | Big sale today | Marker is dropped |
| <$5|five dollars> only | five dollars only | Spoken side is kept |
| || @{reveal!} || | Nothing | 400 transcript_no_speech |
| Plain text | Plain text | Unchanged |
Handling the 400
The error is raised before any job is created, so nothing is billed and nothing is stored. Treat it as a content bug: the script has lines made only of markup. Filter those lines out before you submit, or send them to the caption job instead. A transcript sent through transcript_source keeps its stored text instead, because its receipt is hashed from the source; the worker drops markup from what it sends to the voice.
import re
MARKUP = re.compile(r"\|\||@\{[^{}\n]*\}")
def speakable(line):
return bool(MARKUP.sub(" ", line).strip())
script = ["Hello||world", "|| @{reveal!} ||", "Big sale today"]
print([l for l in script if speakable(l)])
Limits
The Python check above only covers || and @{}; it does not parse <display|spoken> units, which the API handles with its own parser. Use it as a pre-flight filter, not as a copy of the server rule.
Do not rely on this to strip provider tags: a tag such as <break time="1s"/> is not caption markup, and Sume has no SSML field. See the captions docs for how the same wording drives burned-in text.
Where this comes up in practice
The case shows up when a script is written once for a video and reused for narration. The same file then carries timing marks for the caption side and plain words for the voice side. Keeping one source avoids two drifting copies, and it is why the API tolerates the markup instead of rejecting it.
The case to watch is the empty result. A script with a reveal line that is only a marker will pass a length check on your side and still fail on ours, so check each line, not the whole text.
Sources
Related posts
More in Developers
- TTS with only an API key: list avatars and pick one with voice ready
You do not need a voice id for Sume TTS. List your avatars, pick one whose voice status is ready, and send its handle as avatar_handle. Python example.
- TTS word timings to karaoke captions: map words[] to text
Sume TTS returns words[] with start and end; captions take words as text, start, end. A short Python map skips speech-to-text and burns your exact script.
- Turn STT words into paragraphs: break at pauses over 1.2 seconds
Sume STT returns a flat text string. Use word start and end times to break it into paragraphs at long pauses. Python, no extra API call or cost.
- AI music API in TypeScript: generate, wait and save an MP3
Call Sume's Music Router from TypeScript: generateMusicRouter, waitForJob, then download the audio artifact. $0.125 per track, Google lists Lyria 3.5 at $0.08.
Written by Sume