Transcribe a lecture to notes with timestamps
Transcribe a lecture with sentence timestamps, then have a language model turn it into notes whose headings point back to where each topic starts.

To transcribe a lecture to notes, transcribe the recording with sentence timestamps, then have a language model group the sentences into topics and write notes under each, keeping the time each topic starts so you can jump back to that part of the recording. With Sume, STT 1.0 returns the timed sentences and Agent Completions returns the notes as JSON in a shape you define.
The facts come from the STT 1.0 schema in the Sume API reference and from Agent Completions, Structured output, and Usage, all read on 2026-09-28. The two steps are the same as for meeting minutes from a recording; here the output is study notes tied to times.
How do I transcribe a lecture recording?
Send the recording's public HTTPS URL to STT 1.0 with segmentation: {"mode": "sentence"}; the request has no field for file bytes. Sentence mode groups words on terminal punctuation and splits unpunctuated runs on silence, and each segment carries its text with start and end in seconds from the start of the audio. One request covers up to 10 minutes, so split a longer lecture and add each part's start time to its segments, as Transcribe long audio files shows.
If you only need the lecture as text, stop here: the result's text is the whole transcript.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bio101-week3-part-1" \
-d '{
"audio_url": "https://example.com/audio/bio101-week3-part-1.m4a",
"duration_seconds": 600,
"segmentation": { "mode": "sentence" }
}'What should lecture notes contain?
The parts you will study from, in a shape you fix in advance. The example below asks for three: sections, each with a title, the start_seconds where the topic begins, and notes as short lines; key_terms, the vocabulary to learn; and review_questions to test yourself later. Tell the model to start each section at a segment's start and to use only what was said, so the notes stay a record of this lecture rather than the topic in general.
To add a short summary, add a summary string to the schema and to its required list: every property the schema declares must be listed there.
How do I turn the transcript into study notes?
Send the timed segments to POST /v1/agent/completions in input, which the agent treats as data, not instructions, and put the notes' shape in output_schema. The rest works as in the meeting minutes post: the schema follows Sume's strict subset, generation_spend_cap_usd is required, and the call returns a receipt you poll until a completed run fills output.
curl -X POST https://api.sume.com/v1/agent/completions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: bio101-week3-notes" \
-d '{
"instruction": "Turn this lecture into study notes. Start each section at a segment start. Use only what was said.",
"input": { "segments": [{ "index": 0, "text": "Today we cover cell membranes.", "start": 0, "end": 2.6 }] },
"output_schema": { "name": "acme/lecture-notes/v1", "schema": {
"type": "object", "additionalProperties": false,
"required": ["sections", "key_terms", "review_questions"],
"properties": {
"sections": { "type": "array", "items": {
"type": "object", "additionalProperties": false,
"required": ["title", "start_seconds", "notes"],
"properties": {
"title": { "type": "string" }, "start_seconds": { "type": "number" },
"notes": { "type": "array", "items": { "type": "string" } } } } },
"key_terms": { "type": "array", "items": { "type": "string" } },
"review_questions": { "type": "array", "items": { "type": "string" } } } } },
"generation_spend_cap_usd": 1
}'How do I jump from a note back to the recording?
Use each section's start_seconds. Before you rely on it, snap every start_seconds to the nearest real segment start, since the model writes the number, then print it as a clock time next to the heading:
// segments: your time-shifted STT segments; notes: the run's output
const starts = segments.map((s) => s.start);
const snap = (t) =>
starts.reduce((best, s) => (Math.abs(s - t) < Math.abs(best - t) ? s : best));
const clock = (t) => new Date(Math.round(t) * 1000).toISOString().slice(11, 19);
for (const section of notes.sections) {
console.log(`${clock(snap(section.start_seconds))} ${section.title}`);
for (const line of section.notes) console.log(` - ${line}`);
}What does it cost?
Transcription is $0.01 per audio minute, plus a 5.5% agent fee by default. The completion's own cost is debited_usd from GET /v1/usage?run_id=…, the agent's turns included, as in the meeting minutes post.
| Lecture length | STT 1.0 requests | Price before the fee |
|---|---|---|
| 30 minutes | 3 | $0.30 |
| 50 minutes | 5 | $0.50 |
| 90 minutes | 9 | $0.90 |
What are the limits?
- No speaker labels: a question from the audience lands in the transcript without saying who asked it.
- Notes are model output, so check them against the recording. Read
output_errorfirst: if nothing fits your schema,outputisnulland over the API the run endsfailed.
Sources
Related posts
More in Use cases
- Video prospecting with an AI avatar: one clip per prospect
Video prospecting puts a short personal video in a sales outreach message. With an AI avatar, fill one script template per prospect and render each clip.
- What is a video sales letter (VSL)? And making one with AI
A video sales letter (VSL) is a sales pitch delivered as one narrated video, from hook to offer. What goes in one, and how to build it with AI parts.
- Virtual staging AI: furnish an empty room from one photo
Virtual staging with AI: send the empty-room photo with a prompt naming the style and what must not change, then check the walls and windows.
- Virtual try-on for Shopify: live AR or try-on videos
Virtual try-on on Shopify comes two ways: a live camera app on the product page, or AI try-on videos you make ahead and add to the product's media.
Written by Sume