How to create meeting minutes from an audio recording
Create meeting minutes from a recording in two steps: transcribe the audio, then have a language model draft the summary, decisions, and action items.

To create meeting minutes from an audio recording, transcribe the recording to text, then have a language model turn the transcript into minutes: a short summary, the decisions made, and action items with an owner and a due date. Review the draft before you send it: it is model output, written from what the recording says. With Sume, STT 1.0 makes the transcript from a public HTTPS audio URL, and Agent Completions returns the minutes as JSON in a shape you define.
The facts come from the STT 1.0 schema in the Sume API reference and from Agent Completions, Structured output, and Usage, all read on 2026-09-28. For a video already on Sume, Summarize a video with an API does the same with a transcript and stills.
How do I transcribe the meeting recording?
Put the file where STT 1.0 can fetch it: the request takes a public HTTPS audio_url, not an upload. One request covers up to 10 minutes, so split a longer meeting into parts, as Transcribe long audio files shows. Leave out language_code and the language is detected.
- This works on a finished recording, not a live call: the transcription is a job that reads a file at a URL, and the Developer API has no SSE or WebSocket transport today.
- The result's
textis the transcript to pass on. Itswordscarry timings, which you don't need for minutes.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: weekly-sync-part-1" \
-d '{ "audio_url": "https://example.com/audio/weekly-sync-part-1.m4a", "duration_seconds": 600 }'How do I turn the transcript into minutes?
Send it to POST /v1/agent/completions with the task in instruction, the transcript in input, which the agent treats as data and never as instructions, and the minutes' shape in output_schema. generation_spend_cap_usd is required, has no default, and can't be 0. Don't attach the recording itself: completion attachments are images only today.
- For a meeting transcribed in parts, join the parts'
textin order and send it as one transcript. - Sume's strict schema subset has no optional properties: every property is listed in
required, and a field that may be empty allowsnull, asownerandduedo below. Fix output_schema violations covers the other rules. - The API key needs the
agent_completions:writescope, which keys created before Agent Completions shipped don't carry.
curl -X POST https://api.sume.com/v1/agent/completions \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: weekly-sync-minutes" \
-d '{
"instruction": "Write minutes from this meeting transcript. Use only what was said. Leave owner or due null when nobody said one.",
"input": { "transcript": "Okay, first item is the launch date..." },
"output_schema": { "name": "acme/meeting-minutes/v1", "schema": {
"type": "object", "additionalProperties": false,
"required": ["summary", "decisions", "action_items"],
"properties": {
"summary": { "type": "string" },
"decisions": { "type": "array", "items": { "type": "string" } },
"action_items": { "type": "array", "items": {
"type": "object", "additionalProperties": false,
"required": ["task", "owner", "due"],
"properties": {
"task": { "type": "string" },
"owner": { "type": ["string", "null"] },
"due": { "type": ["string", "null"] } } } } } } },
"generation_spend_cap_usd": 1
}'Who owns each action item?
Only people who are named. STT 1.0 returns no speaker labels, so the transcript doesn't say who is talking, and an owner can come only from a name said aloud, as in “Dana will send the budget by Friday.” Allowing null for owner and due, and saying so in the instruction, gives the model a way to leave them empty instead of guessing; fill them in when you review. If each person was recorded on a separate track, How to transcribe an interview shows how to label speakers.
How do I get the minutes back?
The create call returns 202 with an agent.run receipt, not the minutes. Poll its status_url until next_action stops being poll_status; a completed run fills output with your summary, decisions, and action_items. Check output_error first: if nothing satisfied your schema, output is null, and over the API the run ends failed. Run the Sume agent from your backend covers the receipt and webhooks.
What does it cost?
Transcribing a 60-minute meeting takes 6 requests and costs $0.60, plus a 5.5% agent fee by default. The spend cap limits generation, not the agent's own model turns, so read what the completion really cost from its usage summary.
| Step | Call | Price | Limits |
|---|---|---|---|
| Transcribe | POST /v1/stt-1.0/transcribe | $0.01 per audio minute | Up to 10 minutes per request; public HTTPS audio_url |
| Write the minutes | POST /v1/agent/completions | debited_usd from GET /v1/usage?run_id=…, the agent's own turns included | generation_spend_cap_usd required; image attachments only |
Sources
Related posts
More in Use cases
- Create motivational videos with AI: voice, shots, and music
Create motivational videos with AI: a slow spoken quote over cinematic shots, a music bed that builds and ducks under the voice, and big captions.
- Press release video: an announcement read by an AI avatar
A press release video reads the headline, key facts, and a quote in about a minute, captioned from the release text. How to make one with an AI avatar.
- New product launch video with an AI avatar presenter
A new product launch video names the problem, shows what's new, and ends on one call to action. Make it from launch copy with an AI avatar presenter.
- Product URL to video AI: how link-to-video tools work
Product URL to video AI reads a page's title, copy and images, then scripts and renders an ad. How it works, and what to send when a tool can't.
Written by Sume