Pipes and @{} markers in a Sume TTS transcript: stripped, never spoken

Sume strips || cue breaks, @{...} markers and the display side of <display|spoken> before the voice reads; an empty result returns 400 transcript_no_speech.

4 min readSume
All posts

If a TTS transcript contains caption markup, Sume removes it before the voice sees it: a || cue break and an @{...} marker say nothing, and <display|spoken> keeps only the spoken side. The strip runs before validation, billing and storage, so the voice never reads "vertical bar". If nothing speakable is left, for example a transcript of || @{reveal!} ||, the API returns 400 transcript_no_speech with details.field set to transcript.

What gets removed

The rule lives in the API submit path for POST /v1/tts-1.0/generate. A cue break always separates words, so a||b is spoken as two words, and line breaks inside the transcript are kept, because a voice reads them as pauses. Text without any markup comes back byte for byte unchanged.

Captions are the other half of the story: the same markup is meant for caption cue timing, so one script can drive both the voice and the captions without a hand-cleaned copy for each.

How the spoken text is derived (Sume API source, repo read 2026-10-05)
You sendVoice speaksNote
Hello||worldHello worldCue break becomes a space
Big sale @{pop!} todayBig sale todayMarker is dropped
<$5|five dollars> onlyfive dollars onlySpoken side is kept
|| @{reveal!} ||Nothing400 transcript_no_speech
Plain textPlain textUnchanged

Handling the 400

The error is raised before any job is created, so nothing is billed and nothing is stored. Treat it as a content bug: the script has lines made only of markup. Filter those lines out before you submit, or send them to the caption job instead. A transcript sent through transcript_source keeps its stored text instead, because its receipt is hashed from the source; the worker drops markup from what it sends to the voice.

import re

MARKUP = re.compile(r"\|\||@\{[^{}\n]*\}")

def speakable(line):
    return bool(MARKUP.sub(" ", line).strip())

script = ["Hello||world", "|| @{reveal!} ||", "Big sale today"]
print([l for l in script if speakable(l)])

Limits

The Python check above only covers || and @{}; it does not parse <display|spoken> units, which the API handles with its own parser. Use it as a pre-flight filter, not as a copy of the server rule.

Do not rely on this to strip provider tags: a tag such as <break time="1s"/> is not caption markup, and Sume has no SSML field. See the captions docs for how the same wording drives burned-in text.

Where this comes up in practice

The case shows up when a script is written once for a video and reused for narration. The same file then carries timing marks for the caption side and plain words for the voice side. Keeping one source avoids two drifting copies, and it is why the API tolerates the markup instead of rejecting it.

The case to watch is the empty result. A script with a reveal line that is only a marker will pass a length check on your side and still fail on ours, so check each line, not the whole text.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume