Muse Voice Transcribe WebSocket streaming vs Sume STT job waits
Muse Voice Transcribe streams audio over one WebSocket and returns cumulative partials. Sume STT is a job: wait up to 30 seconds, then poll status_url.

Meta's Muse Voice Transcribe is a streaming model: you open one WebSocket, stream audio and read cumulative partial transcripts while the speaker talks. Sume's sume/stt-1.0 is not a stream. You submit a job, optionally wait up to 30 seconds with wait_timeout_seconds, and poll status_url if it is not done.
Meta's protocol is from its developer guide; Sume's from the /v1/stt-1.0/transcribe schema in the API reference, read 2026-10-01.
How does Muse Voice Transcribe deliver text?
The guide describes one WebSocket connection: stream raw audio in, read transcripts back. Partials are cumulative, each one replaces the last, so you render in place, and a final: true frame is the completion signal. The post also describes a one-shot file endpoint: one HTTP POST, one JSON response.
How does a Sume STT job deliver text?
mode is async by default and returns polling URLs right away. sync and subscribe are aliases for a bounded wait of up to wait_timeout_seconds (maximum 30). If the job is still queued or running the response says so and you poll status_url instead of resubmitting. The 30 seconds bounds the HTTP wait, not the job.
| Aspect | Muse Voice Transcribe | Sume `sume/stt-1.0` |
|---|---|---|
| Transport | WebSocket stream, or one HTTP POST for files | HTTPS job: async, sync, subscribe, webhook |
| Partial text while audio arrives | Yes, cumulative partials | No |
| Completion signal | final: true frame | Job status result_ready |
| Longest wait on one call | Not stated | 30 seconds, then poll |
| Audio per request | Not stated for streams | duration_seconds hint max 600 s |
What does the sync call look like?
Ask for a bounded wait, then read status_url if the response says the wait timed out.
curl -X POST https://api.sume.com/v1/stt-1.0/transcribe \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: stt-demo-001" \
-d '{
"audio_url": "https://media.sume.com/artifacts/artf_demo/clip.wav",
"duration_seconds": 120,
"mode": "sync",
"wait_timeout_seconds": 30
}'Is Sume the right fit for live captions?
No. Sume STT has no partials, so a live microphone caption needs a streaming service. It fits recorded clips where you want word timings (always returned, no flag) and a durable job you can poll or receive by webhook. For the same contrast with another vendor see ElevenLabs realtime vs Sume batch STT.
Sources
Related posts
More in Developers
- Nano Banana batch API: Gemini's 24 h batch vs Sume async jobs
Gemini's Batch API trades up to 24 hours of turnaround for higher rate limits. Sume has no batch tier for images: send async or webhook jobs per request.
- Nano Banana Pro 21:9: Gemini's ratio list vs Sume's catalog
Gemini's image docs list 21:9 among ten ratios. Sume's Nano Banana Pro and Nano Banana 2 catalogs include 21:9 too; GPT Image 2.5's list does not.
- Next.js dev MCP endpoint leak: keep your Sume API key out of source
CVE-2026-94486 let a malicious page read source snippets from next dev. Keep the Sume API key in an env var, server-side, and log request ids not headers.
- Next.js 16.3.8 security release: does a Sume webhook route change?
Seven Next.js advisories shipped September 30 in 16.3.8 and 15.5.27. None names Route Handlers, so upgrade, then re-run a signed Sume test delivery.
Written by Sume