Audio transcript to timestamped Markdown notes: Python with Sume STT

Turn a 10-minute recording into Markdown with [mm:ss] sentence lines for Notion or Obsidian using Sume STT sentence segments. Cost is about 10 cents per file.

4 min readSume
All posts

Notes apps have no use for SRT. What a person scanning a transcript wants is one line per sentence with a time they can click back to. Sume STT 1.0 returns exactly the data you need if you ask for sentence segmentation: segments[] with index, text, start, end and duration_seconds, gapless and ordered.

The request

Send segmentation: {"mode": "sentence"} on POST /v1/stt-1.0/transcribe. The audio_url must be public HTTPS, and one job takes at most 10 minutes of audio. Pass duration_seconds (1 to 600) as a reservation hint, otherwise one minute is reserved. Segments are time ranges only. Sume does not slice audio for you here.

Code

import os, time, requests
B = "https://api.sume.com/v1"
H = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}

def run(path, body, key=None):
    h = {**H, **({"Idempotency-Key": key} if key else {})}
    r = requests.post(B + path, json={**body, "mode": "async"}, headers=h)
    r.raise_for_status()
    job = r.json()["data"]["job"]["id"]
    while not requests.get(f"{B}/jobs/{job}/status", headers=H).json()["data"]["terminal"]:
        time.sleep(3)
    res = requests.get(f"{B}/jobs/{job}/result", headers=H)
    res.raise_for_status()
    return res.json()["data"]["result"]

res = run("/stt-1.0/transcribe", {"audio_url": URL, "duration_seconds": 540,
          "segmentation": {"mode": "sentence"}})
lines = []
for s in res["segments"]:
    m, ss = divmod(int(s["start"]), 60)
    lines.append(f"- **{m:02d}:{ss:02d}** {s['text'].strip()}")
open("notes.md", "w").write("\n".join(lines))

Cost and limits

STT 1.0 costs about $0.01 per audio minute. A 9-minute memo is roughly 9 cents, and an hour is about 60 cents across six jobs. For longer recordings, split first. The ten-minute limit guide covers that.

Things the notes will not have

  • No speaker names. Sume fixes diarization off, so every line is unlabeled.
  • No clickable audio links. Obsidian and Notion need your own player link: append #t= plus the start seconds to a media URL if your player supports it.
  • Sentence splitting follows the provider's punctuation. Check one file before you batch.

A header worth adding

Write language_code and language_probability from the result into the note's front matter. A low probability tells you to review before sharing. The language-hint post explains when to pass language_code yourself.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume