Uptime monitor for the Sume API: probe GET /v1/me, spend reads
A 30 second probe of GET /v1/me costs two reads a minute against a 4,800 read Free budget. Python probe and a status table for 401, 429 and 5xx.

An uptime check should prove two things: the API answers, and your key still works. GET /v1/me does both, and it is a read. In the rate limit docs, a read is any GET or HEAD, and reads have their own bucket, forty times the plan's write number on the shipped default. Free is 120 writes and 4,800 reads per minute, so a probe every 30 seconds uses two reads a minute.
Never monitor by submitting a tiny generation. That spends the write budget and wallet, and it tests less than the cheap read does.
Probe
It records the status, latency and the ratelimit-remaining header, and treats only 200 as up.
import json, os, time, urllib.error, urllib.request
def probe(base="https://api.sume.com"):
req = urllib.request.Request(f"{base}/v1/me", headers={"x-api-key": os.environ["SUME_API_KEY"]})
started = time.monotonic()
try:
with urllib.request.urlopen(req, timeout=5) as res:
status, headers = res.status, res.headers
except urllib.error.HTTPError as err:
status, headers = err.code, err.headers
return {
"up": status == 200,
"status": status,
"ms": round((time.monotonic() - started) * 1000),
"reads_left": headers.get("ratelimit-remaining"),
"reset_s": headers.get("ratelimit-reset"),
}
if __name__ == "__main__":
while True:
print(json.dumps(probe()), flush=True) # one read per probe, no write spent
time.sleep(30) # 2 reads a minuteRead the result correctly
| Status | Meaning | Page someone? |
|---|---|---|
| 200 | API and key fine | No |
| 401 | Missing, malformed or revoked key | Yes, but it is a key problem, not an outage |
| 429 | Read budget spent; retry-after says when | Check what else shares the key |
| 5xx or timeout | Service or network trouble | Yes after a few in a row |
Keep it cheap and honest
- Use a key reserved for the monitor so its budget and revocation are separate from production traffic.
- The shipped read multiple is a deployment setting, so
ratelimit-limiton the response is the authority, not the table. - Routes that need no key exist, such as the health route, but an anonymous check does not prove your key works.
- Alert on consecutive failures, not a single one.
Sources
Related posts
More in Developers
- What did one transcription job cost? GET /v1/usage by job_id
Read GET /v1/usage?job_id= and sum data.summary.debited_usd_micros to see what one Sume STT job really cost. Holds and refunds are not counted as spend.
- User closed the tab mid-render: cancel the Sume job or let it finish?
Cancel works only before generation starts; after that you get 409 job_generation_already_started and the job bills. A tested tab-close handler.
- Verify a Sume TTS transcript_receipt SHA-256 yourself in Python
Recompute submitted_transcript_sha256 from your script with NFC and LF canonicalization and compare it to the transcript_receipt on a finished Sume TTS job.
- 4K vertical Short in Timeline: the 2160 cap and 1214x2160
Timeline output width and height top out at 2160 and must be even, so 2160x3840 is refused. What the largest 9:16 frame is and whether a Short needs it.
Written by Sume