Speech to text API in Node.js: transcribe audio with fetch
Transcribe audio in Node.js with built-in fetch: submit to Sume STT, poll the job and print sentence segments with timestamps. 30 lines, no dependencies.

To transcribe an audio file from Node.js, send one authenticated POST to https://api.sume.com/v1/stt-1.0/transcribe with a public HTTPS audio_url, then poll the job until it is completed and read text and words[] from the result. Sume STT 1.0 is $0.01 per audio minute, so a 2 minute clip is two cents. The script is an ES module that uses the built-in fetch, available in current Node releases, with no npm install.
The endpoint is a job API, not a streaming one: you get a job id back and ask for the result when it is ready. That is the right shape for recorded files, and it keeps the client small.
The request and the result
The body needs audio_url. Optional fields are language_code (a hint such as en or ko; omit it for auto-detect), duration_seconds (1 to 600, which lets Sume reserve the right amount; omit it and Sume reserves one minute), segmentation for sentence rows, and metadata, which is stored with your job and never sent to the provider. Send an Idempotency-Key header so a retry does not create a second paid job.
A submit returns 202 with request_id, which is the job id. Poll GET /v1/jobs/{id}/status with a pause between reads, and stop on completed, failed or canceled. Then GET /v1/jobs/{id}/result returns text, language_code, words[] with word, start and end in seconds from the audio start, and segments[] if you asked for them.
| Step | Call | In the script |
|---|---|---|
| Submit | POST /v1/stt-1.0/transcribe | call("POST", ...) with segmentation sentence |
| Wait | GET /v1/jobs/{id}/status | while loop, 2 second pause |
| Read | GET /v1/jobs/{id}/result | one line per segment |
The script
Run SUME_API_KEY=... node stt.mjs https://media.sume.com/artifacts/artf_demo/clip.wav stt-001. Segments print as start-end text.
const [audioUrl, key] = process.argv.slice(2);
const headers = {
Authorization: `Bearer ${process.env.SUME_API_KEY}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
};
async function call(method, path, body) {
const res = await fetch(`https://api.sume.com${path}`, {
method,
headers,
body: body && JSON.stringify(body),
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
return res.json();
}
const job = await call("POST", "/v1/stt-1.0/transcribe", {
audio_url: audioUrl,
duration_seconds: 120,
segmentation: { mode: "sentence" },
});
let status = "";
while (status !== "completed") {
await new Promise((r) => setTimeout(r, 2000));
({ status } = await call("GET", `/v1/jobs/${job.request_id}/status`));
if (["failed", "canceled"].includes(status)) throw new Error(`job ended ${status}`);
}
const result = await call("GET", `/v1/jobs/${job.request_id}/result`);
for (const s of result.segments ?? []) console.log(`${s.start}-${s.end} ${s.text}`);Node.js details worth knowing
Name the file stt.mjs so Node treats it as an ES module and allows top-level await. In a CommonJS project, wrap the body in an async function and call it.
The script asks for sentence segmentation, so the result carries segments[] with start, end and text, and the final loop prints a time-coded line for each. Segments are gapless: each end equals the next start. If you want per-word timing for karaoke captions, read words and skip entries whose type is spacing.
fetch does not reject on an HTTP error, so the helper checks res.ok and throws with the status and body. The Idempotency-Key is the second command-line argument, so a retry from your queue reuses the same job instead of paying twice.
For many files, run several of these at once, but keep an eye on 429 responses and the Retry-After header rather than looping blindly.
Limits and prices
One job takes at most 10 minutes of audio. For longer recordings, cut the audio into slices of up to 600 seconds and send one job per slice, adding each slice's start offset to its word times when you merge. The audio must be at a public HTTPS URL, and Sume media URLs are the preferred source. If the audio is inside a video, audio detach extracts a 16 kHz mono WAV for $0.01 per job.
If the request is refused for balance, the API returns 402; a changed body under a reused idempotency key returns 409; a rate limit returns 429. Do not resubmit a paid request just because your own process timed out, because the job may still be running. Read the status first, as the jobs docs advise.
Keep the returned job id next to your own record of the file. If a result looks wrong later, that id is what lets you fetch the same job again without paying for a second transcription, and the metadata object you sent at submit time is stored with it, so you can tag each job with your own file or ticket id and find it again.
Sources
Related posts
More in Developers
- Speech to text API in Ruby: transcribe audio with Net::HTTP
Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.
- Speech to text API in Swift: transcribe audio with URLSession
Transcribe audio in Swift with async URLSession: submit to Sume STT, poll the job, print text. No packages, 28 lines, one cent per audio minute of audio.
- Spot-check burned-in captions with video frames at STT word times
Check captions on a rendered video by pulling stills at word midpoints from a Sume STT result with POST /v1/video-frames, then compare text to speech.
- Start a Sume render from a serverless function: submit, save, 202
A function must not wait for a video. Submit with mode webhook and a stable Idempotency-Key, save the status URL, return 202, and let the signed webhook finish.
Written by Sume