Speech to text API in Java: transcribe audio with HttpClient

Call a speech to text API from Java with the JDK HttpClient and Jackson: post the audio URL, poll the job, then read the transcript and word times.

6 min readSume
All posts

To use a speech to text API from Java, POST the audio file's URL with the JDK's java.net.http.HttpClient, poll the job until it finishes, then read the transcript and word timestamps from the JSON result. With Sume STT 1.0 that is POST https://api.sume.com/v1/stt-1.0/transcribe, then GET /v1/jobs/{id}/status until terminal is true, then GET /v1/jobs/{id}/result for text, words, and segments.

Sume's SDK is a TypeScript client, so from Java you call the HTTP API directly. The STT facts come from the STT 1.0 schema in the Sume API reference, Jobs and results, and Authentication. The Java calls follow the Java SE 21 docs for HttpClient and HttpRequest.Builder and Jackson's databind README. All were read on 2026-09-28. The same flow in Python is Speech to text in Python.

What do I need before I start?

  • Java 11 or newer: the java.net.http client has been in the JDK since 11.
  • A JSON library. The code uses Jackson's jackson-databind, whose README says to create one ObjectMapper and reuse it.
  • Your API key in the SUME_API_KEY environment variable of a server. Sume's docs say a key is server-side only: never ship it in client JavaScript or a mobile bundle.
  • Send the key as Authorization: Bearer or x-api-key, never both: a request with both is rejected with 401 unauthorized.
  • The audio at a public HTTPS URL. The request takes audio_url and has no field for file bytes.

How do I send the audio from Java?

Build one HttpClient and reuse it; the JDK docs say a built client is immutable and can send multiple requests. The call helper adds the key, sets a 30-second timeout, because a request without one blocks forever, and returns the response's data object. submit posts the body with an Idempotency-Key: a retry with the same key and body returns the original job instead of billing a second one.

import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.URI;
import java.net.http.*;
import java.time.Duration;
import java.util.Map;

public class SumeStt {
  static final HttpClient HTTP = HttpClient.newHttpClient();
  static final ObjectMapper JSON = new ObjectMapper(); // create once, reuse

  static JsonNode call(HttpRequest.Builder req) throws Exception {
    HttpResponse<String> res = HTTP.send(req.timeout(Duration.ofSeconds(30))
        .header("Authorization", "Bearer " + System.getenv("SUME_API_KEY")).build(),
        HttpResponse.BodyHandlers.ofString());
    if (res.statusCode() >= 300) throw new IllegalStateException(res.statusCode() + " " + res.body());
    return JSON.readTree(res.body()).get("data");
  }

  static JsonNode submit(String audioUrl, String idempotencyKey) throws Exception {
    String body = JSON.writeValueAsString(
        Map.of("audio_url", audioUrl, "segmentation", Map.of("mode", "sentence")));
    return call(HttpRequest.newBuilder(URI.create("https://api.sume.com/v1/stt-1.0/transcribe"))
        .header("Content-Type", "application/json").header("Idempotency-Key", idempotencyKey)
        .POST(HttpRequest.BodyPublishers.ofString(body)));
  }

How do I wait for the transcript and read it?

Poll status_url, sleeping for next_poll_after_seconds between reads, until terminal is true; then check that sume_status is completed and read result_url. Don't lean on mode: sync instead: it waits at most 30 seconds, and the API reference says longer waits should submit with async, the default, and poll from your side. /result is only for completed jobs and answers 409 job_not_completed otherwise, and a client-side timeout doesn't cancel the job, so keep the job id to resume from.

In a non-blocking service, HTTP.sendAsync sends the same requests: the JDK docs say it returns a CompletableFuture at once, which completes when the response is available. Schedule each poll after next_poll_after_seconds instead of sleeping a thread.

  public static void main(String[] args) throws Exception {
    JsonNode job = submit("https://example.com/audio/interview.m4a", "interview-001");
    URI statusUrl = URI.create(job.get("status_url").asText());
    JsonNode status = job;
    do {
      Thread.sleep(1000L * Math.max(1, status.path("next_poll_after_seconds").asInt(2)));
      status = call(HttpRequest.newBuilder(statusUrl));
    } while (!status.get("terminal").asBoolean());
    if (!"completed".equals(status.get("sume_status").asText()))
      throw new IllegalStateException("STT job ended as " + status.get("sume_status").asText());

    JsonNode result = call(HttpRequest.newBuilder(URI.create(job.get("result_url").asText())))
        .get("result");
    System.out.println(result.path("language_code").asText() + ": " + result.get("text").asText());
    for (JsonNode w : result.get("words")) {
      if (w.has("start") && !"spacing".equals(w.path("type").asText()))
        System.out.printf("%7.2fs  %s%n", w.get("start").asDouble(), w.get("word").asText());
    }
  }
}

What does the result contain?

In current code a word carries start and end only when the engine returns them, which is why the loop checks has("start"). Jackson's path returns a missing node rather than null for an absent field, so path("type").asText() is an empty string when a word has no type.

From the STT 1.0 result schema in the Sume API reference (API reference docs), read 2026-09-28.
FieldWhat it holds
textThe whole transcript
language_code, language_probabilityThe detected or requested language and the detection confidence, when available
words[]word, start, and end in seconds from the audio start, ordered by start, with an optional type such as word or spacing. Capped at 20,000 entries, flagged by words_truncated.
segments[]Only with segmentation: {"mode": "sentence"}: index, text, start, end, duration_seconds

What are the limits, and what does it cost?

STT 1.0 costs $0.01 per audio minute, plus a 5.5% agent fee by default. If your balance can't cover the estimate, the submit fails with 402 insufficient_credits before any work starts, so call throws on it like any other status of 300 or more.

The limits don't depend on the language you call from: up to 10 minutes of audio per request, no speaker labels, and no live microphone stream, because the job reads a file at a URL. Speech to text in Python lists them in full, and Transcribe long audio files splits longer recordings.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume