Speech to text API in Ruby: transcribe audio with Net::HTTP

Transcribe audio in Ruby with only the standard library: submit to Sume STT, poll the job, print text and word times. A 30-line script at one cent a minute.

5 min readSume
All posts

To transcribe an audio file from Ruby, send one authenticated POST to https://api.sume.com/v1/stt-1.0/transcribe with a public HTTPS audio_url, then poll the job until it is completed and read text and words[] from the result. Sume STT 1.0 is $0.01 per audio minute, so a 2 minute clip is two cents. The script below uses only net/http, json and digest from the standard library, so there is no gem to install.

The endpoint is a job API, not a streaming one: you get a job id back and ask for the result when it is ready. That is the right shape for recorded files, and it keeps the client small.

The request and the result

The body needs audio_url. Optional fields are language_code (a hint such as en or ko; omit it for auto-detect), duration_seconds (1 to 600, which lets Sume reserve the right amount; omit it and Sume reserves one minute), segmentation for sentence rows, and metadata, which is stored with your job and never sent to the provider. Send an Idempotency-Key header so a retry does not create a second paid job.

A submit returns 202 with request_id, which is the job id. Poll GET /v1/jobs/{id}/status with a pause between reads, and stop on completed, failed or canceled. Then GET /v1/jobs/{id}/result returns text, language_code, words[] with word, start and end in seconds from the audio start, and segments[] if you asked for them.

What the Ruby script does at each step, from Sume docs and schema read 2026-10-07
StepCallIn the script
SubmitPOST /v1/stt-1.0/transcribecall(:post, ...) with an Idempotency-Key
WaitGET /v1/jobs/{id}/statusuntil-loop with sleep 2
ReadGET /v1/jobs/{id}/resultputs text and first five words

The script

Run it as ruby stt.rb https://media.sume.com/artifacts/artf_demo/clip.wav after exporting SUME_API_KEY. It stops with a message if the job fails or is canceled.

require "net/http"
require "json"
require "digest"
def call(method, path, body = nil, key = nil)
  uri = URI("https://api.sume.com#{path}")
  req = (method == :post ? Net::HTTP::Post : Net::HTTP::Get).new(uri)
  req["Authorization"] = "Bearer #{ENV.fetch('SUME_API_KEY')}"
  if body
    req["Content-Type"] = "application/json"
    req["Idempotency-Key"] = key
    req.body = JSON.generate(body)
  end
  res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
  raise "HTTP #{res.code}: #{res.body}" unless res.is_a?(Net::HTTPSuccess)
  JSON.parse(res.body)
end
url = ARGV.fetch(0)
job = call(:post, "/v1/stt-1.0/transcribe",
           { audio_url: url, language_code: "en", duration_seconds: 120 },
           "stt-#{Digest::SHA256.hexdigest(url)[0, 16]}")
id = job.fetch("request_id")
status = nil
until %w[completed failed canceled].include?(status)
  sleep 2
  status = call(:get, "/v1/jobs/#{id}/status")["status"]
end
abort "job #{id} ended #{status}" unless status == "completed"
result = call(:get, "/v1/jobs/#{id}/result")
puts result["text"]
result["words"].first(5).each { |w| puts "#{w['start']}-#{w['end']} #{w['word']}" }

Ruby details worth knowing

Ruby's Net::HTTP.start needs use_ssl: true for an HTTPS host, and the block form closes the connection for you. The script checks Net::HTTPSuccess and raises with the response body, which carries Sume's typed error code, so a 402 or 409 explains itself in the exception message.

The idempotency key is a short SHA-256 of the audio URL. The same file produces the same key, so a retry after a crash reuses the first job. If you change language_code or duration_seconds for the same URL, change the key too, because the body is part of the idempotent request.

ENV.fetch('SUME_API_KEY') raises when the variable is missing, which is better than sending a request with an empty bearer token. Keep the key in the environment and out of the script.

Limits and prices

One job takes at most 10 minutes of audio. For longer recordings, cut the audio into slices of up to 600 seconds and send one job per slice, adding each slice's start offset to its word times when you merge. The audio must be at a public HTTPS URL, and Sume media URLs are the preferred source. If the audio is inside a video, audio detach extracts a 16 kHz mono WAV for $0.01 per job.

If the request is refused for balance, the API returns 402; a changed body under a reused idempotency key returns 409; a rate limit returns 429. Do not resubmit a paid request just because your own process timed out, because the job may still be running. Read the status first, as the jobs docs advise.

Keep the returned job id next to your own record of the file. If a result looks wrong later, that id is what lets you fetch the same job again without paying for a second transcription, and the metadata object you sent at submit time is stored with it, so you can tag each job with your own file or ticket id and find it again.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume