Strip Markdown, HTML and URLs before text to speech: Python and cost
Sume TTS bills every character you send, spaces included. Strip Markdown, HTML tags and link URLs in Python first, then see the characters and cents saved.

Strip Markdown, HTML tags and URLs from your text before you send it to Sume text to speech. The API reads transcript as the text to synthesize and bills every character, spaces and punctuation included, at $0.0475 per 1,000. A pasted blog post full of link addresses and ** marks pays for characters you never wanted spoken, and a long address may be read aloud letter by letter. The Python below cleans a post and prints the characters and cost before and after.
The billing rule and the 20,000-character request limit come from API pricing and the Sume API reference, read on 2026-10-03. I also looked through the API code for a step that removes markup before synthesis and found none, so treat the transcript as spoken exactly as sent. I did not generate audio for this post, so how a given voice reads a raw URL is not something I measured.
Why clean the text first?
Two reasons: money and listening. Markup is billed like any other character, and it is the part a listener should never hear. Markdown link syntax is the worst case, because the visible words are short and the address after them is long.
| Markup | Example | Characters billed | What to do |
|---|---|---|---|
| Markdown link | [live now](https://example.com/spring) | Whole string, 38 here | Keep the link text only |
| HTML tag | <p>Free shipping</p> | Tags too, 7 extra here | Drop the tags |
| Bare URL | See https://example.com/faq | Every letter of the address | Remove it or say the page name |
| Emphasis and list marks | **new**, - item, ## Title | Each symbol | Delete the symbols |
How do I strip it in Python?
The function below handles code fences, HTML tags, images, links, bare URLs, list and heading marks and emphasis marks, then adds a full stop to any line that lacks one, so a heading does not run into the next sentence. It uses only the standard library and prints the cost of the raw and the cleaned text.
import re
RATE = 0.0475 / 1000
def clean(text):
text = re.sub(r"```.*?```", " ", text, flags=re.S)
text = re.sub(r"<[^>]+>", " ", text)
text = re.sub(r"!\[[^\]]*\]\([^)]*\)", " ", text)
text = re.sub(r"\[([^\]]+)\]\([^)]*\)", r"\1", text)
text = re.sub(r"https?://\S+", "", text)
text = re.sub(r"^\s{0,3}(#{1,6}|[-*+]|\d+\.)\s+", "", text, flags=re.M)
text = re.sub(r"[*_`>]+", "", text)
lines = [re.sub(r"\s+", " ", ln).strip() for ln in text.splitlines()]
lines = [ln for ln in lines if ln]
return " ".join(ln if ln[-1] in ".!?" else ln + "." for ln in lines)
raw = """## Spring sale
Our **new** line is [live now](https://example.com/spring?utm_source=tts&utm_medium=email).
<p>Free shipping over $50.</p>
- Order by Friday
See https://example.com/faq for details."""
spoken = clean(raw)
print(spoken)
for name, t in (("raw", raw), ("clean", spoken)):
print(name, len(t), "characters =", round(len(t) * RATE, 5), "USD")What does the cleaning save?
On the sample, 196 characters become 96, which is $0.0093 against $0.0046. Pennies on one post, but the saving scales with every link: a 2,000-word article with 25 links carrying 80-character tracking URLs loses about 2,000 characters, roughly $0.10, before a word is spoken. The larger gain is that the take needs no retake because a voice stumbled on an address.
What can go wrong?
Stripping is lossy, so read the output once.
- A removed URL can leave a hole, as in "See for details." in the sample. Rewrite those sentences by hand or replace the address with a spoken page name.
- Prices such as $50 keep their symbol, which is correct. Check how numbers are read with a short test list before a long job.
- Sume takes plain text with no SSML field, so cleaning is the only control you have over what is spoken. The post on SSML in text to speech covers that limit.
- Send
languagewith the cleaned text, and keep the request under 20,000 characters; the character count above is the number the bill uses.
Sources
Related posts
More in Developers
- Stripe branding icon: JPG or PNG under 512 KB, 128 px up
Stripe wants a JPG or PNG icon under 512kb and at least 128 x 128 px. Make a 1:1 mark on Sume, shrink it in Pillow, and know where the icon shows up.
- STT sentence segmentation: unpunctuated speech split on silence
Sume STT 1.0 segmentation mode sentence returns gapless sentence segments, splitting unpunctuated runs on silence. Options, boundary_lead_ms and caveats.
- Submit 20 music takes at once: queued is normal, queue_full is not
Sume accepts music jobs as queued while the queue has room, then returns 429 queue_full. Default accepted-job capacity by plan, from 6 on Free to 120 on Scale.
- subscribeFormatRun onCreated: save the run id, set your own key
subscribeFormatRun makes a new Idempotency-Key per call unless you pass one. Save the run id in onCreated and pass a stable key so a restart cannot double-run.
Written by Sume