Transcribe a lecture to notes with timestamps

Transcribe a lecture with sentence timestamps, then have a language model turn it into notes whose headings point back to where each topic starts.

5 min readSume
All posts

To transcribe a lecture to notes, transcribe the recording with sentence timestamps, then have a language model group the sentences into topics and write notes under each, keeping the time each topic starts so you can jump back to that part of the recording. With Sume, STT 1.0 returns the timed sentences and Agent Completions returns the notes as JSON in a shape you define.

The facts come from the STT 1.0 schema in the Sume API reference and from Agent Completions, Structured output, and Usage, all read on 2026-09-28. The two steps are the same as for meeting minutes from a recording; here the output is study notes tied to times.

How do I transcribe a lecture recording?

Send the recording's public HTTPS URL to STT 1.0 with segmentation: {"mode": "sentence"}; the request has no field for file bytes. Sentence mode groups words on terminal punctuation and splits unpunctuated runs on silence, and each segment carries its text with start and end in seconds from the start of the audio. One request covers up to 10 minutes, so split a longer lecture and add each part's start time to its segments, as Transcribe long audio files shows.

If you only need the lecture as text, stop here: the result's text is the whole transcript.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bio101-week3-part-1" \
  -d '{
    "audio_url": "https://example.com/audio/bio101-week3-part-1.m4a",
    "duration_seconds": 600,
    "segmentation": { "mode": "sentence" }
  }'

What should lecture notes contain?

The parts you will study from, in a shape you fix in advance. The example below asks for three: sections, each with a title, the start_seconds where the topic begins, and notes as short lines; key_terms, the vocabulary to learn; and review_questions to test yourself later. Tell the model to start each section at a segment's start and to use only what was said, so the notes stay a record of this lecture rather than the topic in general.

To add a short summary, add a summary string to the schema and to its required list: every property the schema declares must be listed there.

How do I turn the transcript into study notes?

Send the timed segments to POST /v1/agent/completions in input, which the agent treats as data, not instructions, and put the notes' shape in output_schema. The rest works as in the meeting minutes post: the schema follows Sume's strict subset, generation_spend_cap_usd is required, and the call returns a receipt you poll until a completed run fills output.

curl -X POST https://api.sume.com/v1/agent/completions \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: bio101-week3-notes" \
  -d '{
    "instruction": "Turn this lecture into study notes. Start each section at a segment start. Use only what was said.",
    "input": { "segments": [{ "index": 0, "text": "Today we cover cell membranes.", "start": 0, "end": 2.6 }] },
    "output_schema": { "name": "acme/lecture-notes/v1", "schema": {
      "type": "object", "additionalProperties": false,
      "required": ["sections", "key_terms", "review_questions"],
      "properties": {
        "sections": { "type": "array", "items": {
          "type": "object", "additionalProperties": false,
          "required": ["title", "start_seconds", "notes"],
          "properties": {
            "title": { "type": "string" }, "start_seconds": { "type": "number" },
            "notes": { "type": "array", "items": { "type": "string" } } } } },
        "key_terms": { "type": "array", "items": { "type": "string" } },
        "review_questions": { "type": "array", "items": { "type": "string" } } } } },
    "generation_spend_cap_usd": 1
  }'

How do I jump from a note back to the recording?

Use each section's start_seconds. Before you rely on it, snap every start_seconds to the nearest real segment start, since the model writes the number, then print it as a clock time next to the heading:

// segments: your time-shifted STT segments; notes: the run's output
const starts = segments.map((s) => s.start);
const snap = (t) =>
  starts.reduce((best, s) => (Math.abs(s - t) < Math.abs(best - t) ? s : best));
const clock = (t) => new Date(Math.round(t) * 1000).toISOString().slice(11, 19);

for (const section of notes.sections) {
  console.log(`${clock(snap(section.start_seconds))}  ${section.title}`);
  for (const line of section.notes) console.log(`  - ${line}`);
}

What does it cost?

Transcription is $0.01 per audio minute, plus a 5.5% agent fee by default. The completion's own cost is debited_usd from GET /v1/usage?run_id=…, the agent's turns included, as in the meeting minutes post.

Transcription only, computed from the STT 1.0 rate on API pricing and the 10-minute limit in the Sume API reference (API reference docs), read 2026-09-28.
Lecture lengthSTT 1.0 requestsPrice before the fee
30 minutes3$0.30
50 minutes5$0.50
90 minutes9$0.90

What are the limits?

  • No speaker labels: a question from the audience lands in the transcript without saying who asked it.
  • Notes are model output, so check them against the recording. Read output_error first: if nothing fits your schema, output is null and over the API the run ends failed.

Sources

Related posts

More in Use cases

All Use cases posts

Written by Sume