Voicemail to text for a plumber: 160 messages a month on Sume STT

A plumbing shop with 160 voicemails of about 45 seconds a month pays $1.20 to $1.60 on Sume STT at $0.01 a minute. The bill, the limits and a cost script.

5 min readSume
All posts

A plumbing shop that gets 160 voicemails a month, about 45 seconds each, has 120 minutes of audio to turn into text. On Sume speech-to-text, listed at $0.01 per audio minute, that is $1.20 if usage follows the exact minutes and $1.60 if every message is billed as a whole minute. Either way the month costs less than one sandwich, so the choice is about how messages reach the transcriber, not about price.

The 45-second average is an assumption for the example. Replace it with your own: the script below does the sum for any list of lengths.

What Sume takes and returns

Speech-to-text is POST /v1/stt-1.0/transcribe. The body needs a public HTTPS audio_url, ideally a file already on Sume's media host. A voicemail is far under the 10-minute (600 second) cap on one job. Send duration_seconds when you know it; if you omit it, Sume reserves one minute, which for a 45-second message is about right anyway.

The result holds the text, a words[] array with start and end seconds, and, if you ask for segmentation: {"mode": "sentence"}, sentence segments. Set language_code if you know the caller's language, such as es for a Spanish-speaking customer; leave it out and the language is detected. Every submit needs an Idempotency-Key, and a retried request with the same key does not create a second paid job.

The monthly bill

Sume's rate is per audio minute, so the table shows the two bounds. Microsoft's new MAI-Transcribe-2-Streaming is listed at $0.54 per hour of audio through the end of the year, but it is a streaming model for live transcripts, so it appears only as a point of reference for the same minutes.

Voicemail transcription at 45 seconds each, list rates (read 2026-10-05)
Voicemails a monthAudio minutesSume, exact minutesSume, each rounded upMAI-Transcribe-2-Streaming, same audio
8060$0.60$0.80$0.54
160120$1.20$1.60$1.08
320240$2.40$3.20$2.16

Where the work really is

The model is not the hard part for a small shop; getting the audio is. Phone systems keep voicemail in their own storage, and the transcribe route only reads a public HTTPS file. Check whether your phone provider can send a recording to a URL, or whether you must copy each file to a location that URL can reach. Do not paste a private link behind a login, because the fetch fails.

Once the file is reachable, the flow is the same as any Sume job: submit, poll GET /v1/jobs/{id}/status until it is terminal, then read GET /v1/jobs/{id}/result. The jobs and results page describes the envelope, and the errors and credits page lists 429 rate_limited and queue_full. Back off on both and keep the same idempotency key.

What to do with the text

A transcript is most useful when it changes what happens next. A voicemail that says a basement is flooding should not wait in a list. Because Sume gives you the words with their times, you can match a short list of phrases in your own code and send those messages to a text alert while the rest go to a daily summary. Keep the list short and review misses weekly, since a recognizer can mishear a street name or a part number and a keyword filter will not notice.

  • Search it for an address and a phone number, and put them in the first line of the ticket your dispatcher reads.
  • Flag urgent words such as leak, flood or no heat in your own code. Sume returns the words and times; ranking urgency is your logic, not a Sume feature.
  • Keep the audio too. A transcript is a recognizer's guess, and a name or a street can be wrong, so the dispatcher should be able to listen when it matters.
  • Send the transcript and a link to the audio together, so a person can check the one against the other.

Cost script

The script prices any list of voicemail lengths in seconds under both readings of the per-minute rate, and prints Microsoft's streaming figure for the same audio. It sends nothing.

import math

SUME_PER_MIN = 0.01          # Sume STT list, per audio minute
MAI_STREAM_PER_HOUR = 0.54   # MAI-Transcribe-2-Streaming, through end of 2026

seconds = [45] * 160         # replace with your real message lengths

minutes = sum(seconds) / 60
exact = minutes * SUME_PER_MIN
rounded = sum(math.ceil(s / 60) for s in seconds) * SUME_PER_MIN
mai = sum(seconds) / 3600 * MAI_STREAM_PER_HOUR

print(f"{len(seconds)} voicemails, {minutes:.0f} audio minutes")
print(f"Sume exact minutes: ${exact:.2f}")
print(f"Sume each rounded up: ${rounded:.2f}")
print(f"MAI streaming list: ${mai:.2f}")

Check before you commit

Run a week of real messages and read twenty transcripts against the audio. Microsoft's page says MAI-Transcribe-2-Streaming gives real-time transcripts in 60 languages and is priced through the end of the year (Microsoft AI, read 2026-10-05), which matters if you ever want live captions on a phone call. For stored voicemail, a per-minute file route is the simpler fit. Confirm the live STT rate in the Sume catalog before you budget a year, since list prices can change.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume