Fireflies Talk dictation vs transcribing a recorded file by API

Fireflies Talk is free, unlimited desktop dictation. A recorded file needs a different tool. How to tell which you need, and what Sume STT costs for the file.

4 min readSume
All posts

Fireflies Talk dictates live into whatever app you are typing in; a recorded interview or call needs a file transcribed instead. Fireflies launched Talk on October 1 as a free, unlimited voice dictation tool for Mac and Windows with 100+ languages, according to Fossbytes. If you speak into a text box, that fits. If you already have an MP3, you need an upload-and-transcribe path, such as Sume STT at about one cent a minute.

The two tools solve different jobs, so the right question is whether your audio exists yet.

Two jobs, side by side

Only the Fireflies facts come from Fossbytes. Sume facts come from the API schema in the repository.

Live dictation versus file transcription (read 2026-10-08)
QuestionFireflies TalkSume STT
InputYour live voice, in any desktop appA public HTTPS audio_url
PriceFree and unlimited at launchAbout $0.01 per minute
Languages100+Auto-detect, or a language_code hint
PlatformsMac and WindowsAny client that can call the API
OutputText typed where your cursor isWord-timed transcript from a job
Longest inputNot stated10 minutes per request

Transcribe a recorded file in four steps

  • Upload the file so it has a public HTTPS address; a Sume media URL is preferred.
  • POST it to /v1/stt-1.0/transcribe with an Idempotency-Key and Bearer key.
  • Send duration_seconds (1 to 600) so the reservation matches the clip; leaving it out reserves one minute.
  • Read job.id, status_url and result_url from the response, then poll /v1/jobs/:id/status and fetch /v1/jobs/:id/result.
import os, requests

r = requests.post(
    "https://api.sume.com/v1/stt-1.0/transcribe",
    headers={
        "Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Idempotency-Key": "stt-demo-001",
    },
    json={
        "audio_url": "https://media.sume.com/example/clip.mp3",
        "duration_seconds": 540,
    },
    timeout=30,
)
r.raise_for_status()
print(r.json())

When live dictation is the wrong tool

Dictation tools assume one speaker, close to a microphone, who talks and sees the text appear. Meetings, phone calls and field recordings break each of those assumptions. A file transcript lets you re-run the same audio with another tool, compare outputs, and keep an original for review, which a live session does not.

It also separates capture from cost. You can record now and decide later how much transcription to buy. At about one cent per minute, an hour of recorded audio is roughly 60 cents on Sume, so the cost question is often smaller than the workflow question. If you only need a few spoken notes a day, a free dictation tool may be all you ever use, and that is a fine outcome.

What Sume does not do

Sume does not dictate into other apps and has no desktop client for live voice. It does not label speakers. Anything longer than 10 minutes must be cut into pieces before upload. Sume STT does not run on your device, does not stream, does not label speakers (diarization and audio-event tagging are fixed off on the server), and takes at most 10 minutes of audio per request.

Which one

If the words are in your head, dictate. If the words are in a file, transcribe it. For a 100-file batch, see the cost table of ten-minute clips. For privacy details on both, read where dictation audio goes.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume