Speech to text API in Ruby: transcribe audio with Net::HTTP
Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.

To transcribe an audio file from Ruby, send one authenticated POST to https://api.sume.com/v1/stt-1.0/transcribe with a public HTTPS audio_url, then poll the job until it is completed and read text and words[] from the result. Sume STT 1.0 is $0.01 per audio minute, so a 2 minute clip is two cents. The script below uses only net/http, json and digest from the standard library, so there is no gem to install.
The endpoint is a job API, not a streaming one: you get a job id back and ask for the result when it is ready. That is the right shape for recorded files, and it keeps the client small.
The request and the result
The body needs audio_url. Optional fields are language_code (a hint such as en or ko; omit it for auto-detect), duration_seconds (1 to 600, which lets Sume reserve the right amount; omit it and Sume reserves one minute), segmentation for sentence rows, and metadata, which is stored with your job and never sent to the provider. Send an Idempotency-Key header so a retry does not create a second paid job.
A submit returns 202 with request_id, which is the job id. Poll GET /v1/jobs/{id}/status with a pause between reads, and stop on completed, failed or canceled. Then GET /v1/jobs/{id}/result returns text, language_code, words[] with word, start and end in seconds from the audio start, and segments[] if you asked for them.
| Step | Call | In the script |
|---|---|---|
| Submit | POST /v1/stt-1.0/transcribe | call(:post, ...) with an Idempotency-Key |
| Wait | GET /v1/jobs/{id}/status | until-loop with sleep 2 |
| Read | GET /v1/jobs/{id}/result | puts text and first five words |
The script
Run it as ruby stt.rb https://media.sume.com/artifacts/artf_demo/clip.wav after exporting SUME_API_KEY. It stops with a message if the job fails or is canceled.
require "net/http"
require "json"
require "digest"
def call(method, path, body = nil, key = nil)
uri = URI("https://api.sume.com#{path}")
req = (method == :post ? Net::HTTP::Post : Net::HTTP::Get).new(uri)
req["Authorization"] = "Bearer #{ENV.fetch('SUME_API_KEY')}"
if body
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = JSON.generate(body)
end
res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
raise "HTTP #{res.code}: #{res.body}" unless res.is_a?(Net::HTTPSuccess)
JSON.parse(res.body)
end
url = ARGV.fetch(0)
job = call(:post, "/v1/stt-1.0/transcribe",
{ audio_url: url, language_code: "en", duration_seconds: 120 },
"stt-#{Digest::SHA256.hexdigest(url)[0, 16]}")
id = job.fetch("request_id")
status = nil
until %w[completed failed canceled].include?(status)
sleep 2
status = call(:get, "/v1/jobs/#{id}/status")["status"]
end
abort "job #{id} ended #{status}" unless status == "completed"
result = call(:get, "/v1/jobs/#{id}/result")
puts result["text"]
result["words"].first(5).each { |w| puts "#{w['start']}-#{w['end']} #{w['word']}" }Ruby details worth knowing
Ruby's Net::HTTP.start needs use_ssl: true for an HTTPS host, and the block form closes the connection for you. The script checks Net::HTTPSuccess and raises with the response body, which carries Sume's typed error code, so a 402 or 409 explains itself in the exception message.
The idempotency key is a short SHA-256 of the audio URL. The same file produces the same key, so a retry after a crash reuses the first job. If you change language_code or duration_seconds for the same URL, change the key too, because the body is part of the idempotent request.
ENV.fetch('SUME_API_KEY') raises when the variable is missing, which is better than sending a request with an empty bearer token. Keep the key in the environment and out of the script.
Limits and prices
One job takes at most 10 minutes of audio. For longer recordings, cut the audio into slices of up to 600 seconds and send one job per slice, adding each slice's start offset to its word times when you merge. The audio must be at a public HTTPS URL, and Sume media URLs are the preferred source. If the audio is inside a video, audio detach extracts a 16 kHz mono WAV for $0.01 per job.
If the request is refused for balance, the API returns 402; a changed body under a reused idempotency key returns 409; a rate limit returns 429. Do not resubmit a paid request just because your own process timed out, because the job may still be running. Read the status first, as the jobs docs advise.
Keep the returned job id next to your own record of the file. If a result looks wrong later, that id is what lets you fetch the same job again without paying for a second transcription, and the metadata object you sent at submit time is stored with it, so you can tag each job with your own file or ticket id and find it again.
Sources
Related posts
More in Developers
- Speech to text API in Swift: transcribe audio with URLSession
Transcribe audio in Swift with async URLSession: submit to Sume STT, poll the job, print text. No packages, 28 lines, one cent per audio minute of audio.
- Spot-check burned-in captions with video frames at STT word times
Check captions on a rendered video by pulling stills at word midpoints from a Sume STT result with POST /v1/video-frames, then compare text to speech.
- Start a Sume render from a serverless function: submit, save, 202
A function must not wait for a video. Submit with mode webhook and a stable Idempotency-Key, save the status URL, return 202, and let the signed webhook finish.
- A Sume job looks stuck: wait, cancel or poll the events?
Read status, then events. Queued and processing mean wait, cancel only works before generation starts, and a client timeout never cancels the job.
Written by Sume