Transcribe an iPhone voice memo (m4a) with an API: audio_url, limits

Sume STT takes a public HTTPS audio_url up to 10 minutes. What is verified for m4a, why media-imports cannot host a memo, and a Python example.

4 min readSume
All posts

To transcribe an iPhone voice memo with Sume STT, put the .m4a at a public HTTPS URL and call POST /v1/stt-1.0/transcribe with it as audio_url. The OpenAPI examples use .m4a URLs, and the request validation checks the URL, not the file format. Each request covers at most 600 seconds (10 minutes) and is priced at $0.01 per audio minute.

Read on 2026-10-03: the Sume API reference for STT 1.0, the media-imports schema, the Media inputs docs page and the pricing catalog in the Sume repository.

What must audio_url be, and is m4a supported?

Verified in the request schema: a string of at most 2,048 characters, HTTPS, no username or password, no port and no private or local host. The field description says to prefer a Sume media or attachment URL. The published STT examples include https://media.sume.com/artifacts/example/clip.m4a, so m4a is a shape the API documents.

What I could not verify: a list of accepted codecs. The check stops at the URL, so a format problem would surface as a failed job. If an m4a fails, convert it first and send the converted file, for example ffmpeg -i memo.m4a -ac 1 -ar 16000 memo.wav.

Can media-imports host the memo?

No. POST /v1/media-imports accepts public TikTok or Instagram video or reel URLs only (YouTube and other hosts return unsupported_platform), at $0.15 per import. There is no documented upload endpoint for a local voice memo, so host the file yourself on any storage that serves it over public HTTPS.

STT 1.0 request limits, read 2026-10-03
FieldRule
audio_urlPublic HTTPS, up to 2,048 characters
duration_secondsOptional integer 1 to 600; omitted reserves 1 minute
language_codeOptional hint such as en; omit to auto-detect
Resulttext plus words[] with word, start, end in seconds

Runnable example

Pass the URL and the memo length in seconds, so the reservation matches the audio.

import hashlib, os, sys, time, requests

API = "https://api.sume.com"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
url, seconds = sys.argv[1], int(sys.argv[2])
if seconds > 600:
    sys.exit("STT takes at most 600 s per request; split the memo first")
key = hashlib.sha256(f"{url}{seconds}".encode()).hexdigest()[:32]
r = requests.post(f"{API}/v1/stt-1.0/transcribe",
                  headers={**AUTH, "Idempotency-Key": key},
                  json={"audio_url": url, "duration_seconds": seconds})
r.raise_for_status()
job = r.json()["data"]
while True:
    s = requests.get(job["status_url"], headers=AUTH).json()["data"]
    if s["terminal"]:
        break
    time.sleep(s.get("next_poll_after_seconds") or 3)
res = requests.get(job["result_url"], headers=AUTH).json()["data"]["result"]
print(res["text"])

What if the memo is longer than 10 minutes?

Split it into parts of 10 minutes or less, transcribe each, and add each part's start time to its word timestamps. A 25-minute memo is three requests.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume