Dictate a video brief by voice in Sume Agents: the 2-minute limit

The Sume Agents composer records up to 120 seconds from your mic, transcribes it with Sume STT, and adds the text. Here is what happens to the audio.

5 min readSume
All posts

Yes: the Sume Agents composer has a mic button. You record up to 120 seconds, the browser converts the audio if needed, Sume STT transcribes it, and the transcript lands in the composer as text you can edit before sending. Longer briefs should be dictated in two takes, since each recording stops at two minutes.

What the recorder does

The recording limit is a constant in the web app, 120 seconds. The recorder picks the first MIME type the browser supports from audio/webm;codecs=opus, audio/webm, audio/mp4 and audio/mpeg, so the container depends on your browser.

What happens to the file

If the container is not one the transcription route takes directly, such as WebM, the browser decodes it and re-encodes it as 16-bit PCM WAV before upload. The WAV is uploaded through a presigned route, then the app calls its transcription endpoint and inserts the returned text into the composer.

The result is text, not audio. The agent never hears your voice; it reads the transcript you sent, the same as if you had typed it.

What it costs

The transcription runs on Sume STT 1.0, which is listed at $0.01 per audio minute. A full 120-second dictation is two minutes of audio. Your usage page shows what a workspace is charged, since the composer path is a product surface rather than an API call you make.

Writing a brief you can speak

Speech runs long and loose, so give yourself a shape before you press record.

  • Say the product and the one-sentence promise first.
  • Name the length and the aspect ratio: 15 seconds vertical, or 30 seconds square.
  • Say how it should sound: voice style and music mood in one phrase each.
  • End with the single action you want the viewer to take.
  • Read the transcript before sending. Names and product terms are the usual misses; fix them in the box.

When to use the API instead

If you want a recording longer than two minutes, or word timings for captions, call Sume STT directly. It takes a public HTTPS audio_url of up to 10 minutes and returns words[]. For speech pulled from a video, run audio detach first.

If the transcript looks wrong

Edit the text in the composer before you send; nothing is acted on until you do. If a whole recording failed to transcribe, record it again in a quieter spot, shorter, and watch for the 120-second stop. A brief that rambles produces a rambling plan, so trim the transcript to the facts the agent needs.

Browser permission for the microphone is required the first time; if you denied it, re-allow it in the site settings.

Related posts

More in Agents

All Agents posts

Written by Sume