Transcribe an iPhone voice memo (m4a) with an API: audio_url, limits
Sume STT takes a public HTTPS audio_url up to 10 minutes. What is verified for m4a, why media-imports cannot host a memo, and a Python example.

To transcribe an iPhone voice memo with Sume STT, put the .m4a at a public HTTPS URL and call POST /v1/stt-1.0/transcribe with it as audio_url. The OpenAPI examples use .m4a URLs, and the request validation checks the URL, not the file format. Each request covers at most 600 seconds (10 minutes) and is priced at $0.01 per audio minute.
Read on 2026-10-03: the Sume API reference for STT 1.0, the media-imports schema, the Media inputs docs page and the pricing catalog in the Sume repository.
What must audio_url be, and is m4a supported?
Verified in the request schema: a string of at most 2,048 characters, HTTPS, no username or password, no port and no private or local host. The field description says to prefer a Sume media or attachment URL. The published STT examples include https://media.sume.com/artifacts/example/clip.m4a, so m4a is a shape the API documents.
What I could not verify: a list of accepted codecs. The check stops at the URL, so a format problem would surface as a failed job. If an m4a fails, convert it first and send the converted file, for example ffmpeg -i memo.m4a -ac 1 -ar 16000 memo.wav.
Can media-imports host the memo?
No. POST /v1/media-imports accepts public TikTok or Instagram video or reel URLs only (YouTube and other hosts return unsupported_platform), at $0.15 per import. There is no documented upload endpoint for a local voice memo, so host the file yourself on any storage that serves it over public HTTPS.
| Field | Rule |
|---|---|
| audio_url | Public HTTPS, up to 2,048 characters |
| duration_seconds | Optional integer 1 to 600; omitted reserves 1 minute |
| language_code | Optional hint such as en; omit to auto-detect |
| Result | text plus words[] with word, start, end in seconds |
Runnable example
Pass the URL and the memo length in seconds, so the reservation matches the audio.
import hashlib, os, sys, time, requests
API = "https://api.sume.com"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
url, seconds = sys.argv[1], int(sys.argv[2])
if seconds > 600:
sys.exit("STT takes at most 600 s per request; split the memo first")
key = hashlib.sha256(f"{url}{seconds}".encode()).hexdigest()[:32]
r = requests.post(f"{API}/v1/stt-1.0/transcribe",
headers={**AUTH, "Idempotency-Key": key},
json={"audio_url": url, "duration_seconds": seconds})
r.raise_for_status()
job = r.json()["data"]
while True:
s = requests.get(job["status_url"], headers=AUTH).json()["data"]
if s["terminal"]:
break
time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(job["result_url"], headers=AUTH).json()["data"]["result"]
print(res["text"])What if the memo is longer than 10 minutes?
Split it into parts of 10 minutes or less, transcribe each, and add each part's start time to its word timestamps. A 25-minute memo is three requests.
Sources
Related posts
More in Developers
- Translate an SRT and burn it in: Sume caption cues, limits, Python
Sume takes no SRT upload, but caption cues take the same text and times. A Python converter, the 200-cue and 60-second limits, and which fonts apply.
- Trigger.dev 4.6.3 error cause chains: log the Sume code and request id
Trigger.dev v4.6.3 shows thrown error cause chains in the dashboard, CLI and alerts. Wrap a failed Sume call so the cause carries error.code and request_id.
- Trim silence from a TTS file: word timestamps and Timeline audio
Cut the lead-in and tail of a Sume TTS file by reading words[] start and end, then slice it with Timeline audio split for $0.01 or use source_in in a render.
- Try-on double click: one Idempotency-Key per shopper and garment
Stop a double-clicked Try on button from running a Sume Format twice. Derive the Idempotency-Key from shopper, garment and version, and read the 200 replay.
Written by Sume