Text to speech API in Java: pick a voice, save the MP3

Call a text to speech API from Java: choose a voice selector and an audio format, POST the text, wait for the job, then write the MP3 to disk.

5 min readSume
All posts

To use a text to speech API from Java, POST the text and a voice, poll the job until it finishes, then download the audio file the result links to. With Sume TTS 1.0 that is POST https://api.sume.com/v1/tts-1.0/generate, then GET /v1/jobs/{id}/status until terminal is true, then GET /v1/jobs/{id}/result for the MP3's audio_url.

Sume's facts come from the TTS 1.0 schema in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results; the code follows the Java SE 21 BodyHandlers page for the file download. All were read on 2026-09-28. The project setup and the shared request helper are explained line by line in Speech to text API in Java; this page covers what is different when the output is audio.

What goes in the request?

A JSON body with the text and exactly one way to pick the voice: an avatar's voice, or a voice id you already hold. Everything else is optional; the audio defaults to MP3. Keep the key in a server-side environment variable, as Sume's authentication docs say.

From the TTS 1.0 request schema in the Sume API reference, read 2026-09-28.
Body fieldEffect on the audio you get back
transcriptThe text, 1–20,000 characters; spaces and punctuation count toward usage
avatar_handle or avatar_idUses the voice of an avatar from GET /v1/avatar-1.0/avatars whose voice.status is ready
voice.idA TTS voice UUID or a Voices library id (voi_ + 32 hex) you already hold
languageBCP-47 / ISO-639 code; set it for every non-English transcript
output_formatDefaults to mp3, 44100 Hz, 128 kbps; wav and raw are the other containers
timestamps.wordsWhen true, the completed result includes words[] with start/end seconds

What does the Java request helper look like?

The same call helper as the speech to text sample: it adds the bearer key and a 30-second timeout to each request, throws on a non-2xx status, and returns the JSON data object. Two JDK pieces are specific to sending text and getting a file back. BodyPublishers.ofString converts the JSON body to bytes as UTF-8, so accented or non-Latin characters in the transcript arrive intact. BodyHandlers.ofFile(Path) returns an HttpResponse<Path>: the body goes to the file, and the response hands back its path:

import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.net.URI;
import java.net.http.*;
import java.nio.file.Path;
import java.time.Duration;
import java.util.Map;

public class SumeTts {
  static final HttpClient HTTP = HttpClient.newHttpClient();
  static final ObjectMapper JSON = new ObjectMapper(); // create once, reuse

  static JsonNode call(HttpRequest.Builder req) throws Exception {
    HttpResponse<String> res = HTTP.send(req.timeout(Duration.ofSeconds(30))
        .header("Authorization", "Bearer " + System.getenv("SUME_API_KEY")).build(),
        HttpResponse.BodyHandlers.ofString());
    if (res.statusCode() >= 300) throw new IllegalStateException(res.statusCode() + " " + res.body());
    return JSON.readTree(res.body()).get("data");
  }

How do I wait for the audio and save it?

Submit with an Idempotency-Key, so a retried POST returns the original job instead of billing a second one. Poll status_url, sleeping for next_poll_after_seconds, until terminal is true, and read result_url only when sume_status is completed: /result answers 409 job_not_completed until result_ready is true. In current code audio_url is the audio artifact's public Sume CDN URL, so the MP3 download sends no key, and BodyHandlers.ofFile writes the bytes to speech.mp3 on disk:

  public static void main(String[] args) throws Exception {
    String body = JSON.writeValueAsString(Map.of(
        "transcript", "Your order has shipped and arrives on Thursday.", "avatar_handle", "acme"));
    JsonNode job = call(HttpRequest.newBuilder(URI.create("https://api.sume.com/v1/tts-1.0/generate"))
        .header("Content-Type", "application/json").header("Idempotency-Key", "order-shipped-001")
        .POST(HttpRequest.BodyPublishers.ofString(body)));
    JsonNode status = job;
    do {
      Thread.sleep(1000L * Math.max(1, status.path("next_poll_after_seconds").asInt(2)));
      status = call(HttpRequest.newBuilder(URI.create(job.get("status_url").asText())));
    } while (!status.get("terminal").asBoolean());
    if (!"completed".equals(status.get("sume_status").asText()))
      throw new IllegalStateException("TTS job ended as " + status.get("sume_status").asText());

    JsonNode result = call(HttpRequest.newBuilder(URI.create(job.get("result_url").asText())))
        .get("result");
    HttpResponse<Path> audio = HTTP.send(HttpRequest.newBuilder(URI.create(result.get("audio_url").asText()))
        .timeout(Duration.ofSeconds(60)).build(), HttpResponse.BodyHandlers.ofFile(Path.of("speech.mp3")));
    if (audio.statusCode() != 200) throw new IllegalStateException("download " + audio.statusCode());
  }
}

Can I get the speech back in one call or as a stream?

Not as a stream: the API reference describes TTS 1.0 as an async job with poll or webhook, non-streaming, so you get a finished audio file, not chunks to play while they arrive. mode: "sync" holds the submit for at most 30 seconds, and that ceiling bounds the HTTP wait, not the job, so keep the poll loop for long transcripts. A client-side timeout doesn't cancel the job, which keeps running and still bills.

What does text to speech from Java cost?

TTS 1.0 costs $0.0475 per 1,000 characters, plus a 5.5% agent fee by default, and the language you call from doesn't change it. Synthesized audio longer than 1,200 seconds fails with tts_duration_exceeded and captures no credits. WAV output, word timings and per-sentence clips are body changes, listed in Text to speech API in Python.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume