Muse Voice Transcribe WebSocket streaming vs Sume STT job waits

Muse Voice Transcribe streams audio over one WebSocket and returns cumulative partials. Sume STT is a job: wait up to 30 seconds, then poll status_url.

4 min readSume
All posts

Meta's Muse Voice Transcribe is a streaming model: you open one WebSocket, stream audio and read cumulative partial transcripts while the speaker talks. Sume's sume/stt-1.0 is not a stream. You submit a job, optionally wait up to 30 seconds with wait_timeout_seconds, and poll status_url if it is not done.

Meta's protocol is from its developer guide; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01.

How does Muse Voice Transcribe deliver text?

The guide describes one WebSocket connection: stream raw audio in, read transcripts back. Partials are cumulative, each one replaces the last, so you render in place, and a final: true frame is the completion signal. The post also describes a one-shot file endpoint: one HTTP POST, one JSON response.

How does a Sume STT job deliver text?

mode is async by default and returns polling URLs right away. sync and subscribe are aliases for a bounded wait of up to wait_timeout_seconds (maximum 30). If the job is still queued or running the response says so and you poll status_url instead of resubmitting. The 30 seconds bounds the HTTP wait, not the job.

Delivery model, Meta guide vs Sume STT schema, read 2026-10-01.
AspectMuse Voice TranscribeSume `sume/stt-1.0`
TransportWebSocket stream, or one HTTP POST for filesHTTPS job: async, sync, subscribe, webhook
Partial text while audio arrivesYes, cumulative partialsNo
Completion signalfinal: true frameJob status result_ready
Longest wait on one callNot stated30 seconds, then poll
Audio per requestNot stated for streamsduration_seconds hint max 600 s

What does the sync call look like?

Ask for a bounded wait, then read status_url if the response says the wait timed out.

curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: stt-demo-001" \
  -d '{
    "audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
    "duration_seconds": 120,
    "mode": "sync",
    "wait_timeout_seconds": 30
  }'

Is Sume the right fit for live captions?

No. Sume STT has no partials, so a live microphone caption needs a streaming service. It fits recorded clips where you want word timings (always returned, no flag) and a durable job you can poll or receive by webhook. For the same contrast with another vendor see ElevenLabs realtime vs Sume batch STT.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume