Dictate a video brief by voice in Sume Agents: the 2-minute limit
The Sume Agents composer records up to 120 seconds from your mic, transcribes it with Sume STT, and adds the text. Here is what happens to the audio.

Yes: the Sume Agents composer has a mic button. You record up to 120 seconds, the browser converts the audio if needed, Sume STT transcribes it, and the transcript lands in the composer as text you can edit before sending. Longer briefs should be dictated in two takes, since each recording stops at two minutes.
What the recorder does
The recording limit is a constant in the web app, 120 seconds. The recorder picks the first MIME type the browser supports from audio/webm;codecs=opus, audio/webm, audio/mp4 and audio/mpeg, so the container depends on your browser.
What happens to the file
If the container is not one the transcription route takes directly, such as WebM, the browser decodes it and re-encodes it as 16-bit PCM WAV before upload. The WAV is uploaded through a presigned route, then the app calls its transcription endpoint and inserts the returned text into the composer.
The result is text, not audio. The agent never hears your voice; it reads the transcript you sent, the same as if you had typed it.
What it costs
The transcription runs on Sume STT 1.0, which is listed at $0.01 per audio minute. A full 120-second dictation is two minutes of audio. Your usage page shows what a workspace is charged, since the composer path is a product surface rather than an API call you make.
Writing a brief you can speak
Speech runs long and loose, so give yourself a shape before you press record.
- Say the product and the one-sentence promise first.
- Name the length and the aspect ratio: 15 seconds vertical, or 30 seconds square.
- Say how it should sound: voice style and music mood in one phrase each.
- End with the single action you want the viewer to take.
- Read the transcript before sending. Names and product terms are the usual misses; fix them in the box.
When to use the API instead
If you want a recording longer than two minutes, or word timings for captions, call Sume STT directly. It takes a public HTTPS audio_url of up to 10 minutes and returns words[]. For speech pulled from a video, run audio detach first.
If the transcript looks wrong
Edit the text in the composer before you send; nothing is acted on until you do. If a whole recording failed to transcribe, record it again in a quieter spot, shorter, and watch for the 120-second stop. A brief that rambles produces a rambling plan, so trim the transcript to the facts the agent needs.
Browser permission for the microphone is required the first time; if you denied it, re-allow it in the site settings.
Related posts
More in Agents
- Does the Sume run spend cap include the LLM turn? No, only generation
On a Sume run receipt, billable_amount_usd_micros is generation spend only; the agent's LLM turn bills a separate Agent wallet. How to budget both.
- Fire a Sume schedule from a decision, not a clock: API trigger
Set api_trigger_enabled on a Sume schedule, let Clef or Decider decide when it is worth running, then POST /v1/actions/{action_id}/runs with an idempotency key.
- Five generate_image calls in one turn: Sume MCP create budget
Hosted Sume MCP gives paid creates their own budget (20 per principal, 64 per process) and queues a call up to 20 seconds before wait_busy.
- usage.cap on a Format run: an agent reads remaining_usd_micros
A Format run receipt splits its spend cap into limit, counted and remaining USD micros. A supervising agent can read headroom before it asks for more work.
Written by Sume