C# speech to text: transcribe audio files with HttpClient

Speech to text in C#: POST the audio file's URL with HttpClient, poll the job, then read the transcript and word timestamps from the JSON result.

5 min readSume
All posts

To do speech to text in C#, send the audio file's URL to a speech to text API with HttpClient, poll the job until it finishes, then read the transcript and word timestamps from the JSON result. With Sume STT 1.0 that is POST https://api.sume.com/v1/stt-1.0/transcribe with an audio_url, then GET status_url until terminal is true, then GET result_url for text and words[], at up to 10 minutes of audio per request.

Sume's SDK is a TypeScript client, so from C# or an ASP.NET app you call the HTTP API directly. The Sume facts come from the STT 1.0 schema in the Sume API reference (the OpenAPI document behind the API reference docs) and Jobs and results; the .NET calls follow Microsoft's HttpClient and HTTP requests pages. All were read on 2026-09-29. The same flow in Java is Speech to text API in Java.

What do I need before I start?

  • The recording at a public HTTPS URL. The request takes audio_url and has no field for file bytes, and Sume has no public route for uploading local files.
  • At most 10 minutes of audio per request. Longer files go in parts: Transcribe long audio files shows the split.
  • Your API key in a server-side environment variable, sent as Authorization: Bearer. Microsoft's docs name IHttpClientFactory as an alternative to one shared HttpClient.

How do I send the audio from C#?

Reuse one HttpClient, put the key in DefaultRequestHeaders, and send an Idempotency-Key so a retried submit returns the original job instead of billing a second one. duration_seconds (1–600) sizes the usage reservation; without it, 1 minute is reserved. segmentation with mode sentence adds sentence segments to the result.

using System.Net.Http.Headers;
using System.Net.Http.Json;
using System.Text.Json;

var api = new HttpClient(); // create once, reuse
api.DefaultRequestHeaders.Authorization = new AuthenticationHeaderValue(
    "Bearer", Environment.GetEnvironmentVariable("SUME_API_KEY"));

async Task<JsonElement> Data(HttpResponseMessage res)
{
    res.EnsureSuccessStatusCode(); // throws outside 200-299
    return (await res.Content.ReadFromJsonAsync<JsonElement>()).GetProperty("data");
}

var submit = new HttpRequestMessage(HttpMethod.Post, "https://api.sume.com/v1/stt-1.0/transcribe")
{
    Content = JsonContent.Create(new Dictionary<string, object>
    {
        ["audio_url"] = "https://example.com/audio/interview.m4a",
        ["duration_seconds"] = 420,
        ["segmentation"] = new Dictionary<string, string> { ["mode"] = "sentence" },
    }),
};
submit.Headers.Add("Idempotency-Key", "interview-001");
var job = await Data(await api.SendAsync(submit));

How do I read the transcript and word times?

Poll status_url at least once (the submit envelope has no sume_status, and a replayed submit can return a finished job) until terminal is true, then read result_url only if sume_status is completed; /result answers 409 job_not_completed until result_ready is true. In current code a word carries start and end only when the engine returns them, so the loop checks for start.

var status = job; // poll at least once: a replayed submit may already be done
do
{
    var hint = status.GetProperty("next_poll_after_seconds");
    await Task.Delay(TimeSpan.FromSeconds(hint.ValueKind == JsonValueKind.Number ? hint.GetInt32() : 2));
    status = await Data(await api.GetAsync(job.GetProperty("status_url").GetString()));
} while (!status.GetProperty("terminal").GetBoolean());
var state = status.GetProperty("sume_status").GetString();
if (state != "completed") throw new Exception($"STT job ended as {state}");

var result = (await Data(await api.GetAsync(job.GetProperty("result_url").GetString())))
    .GetProperty("result");
Console.WriteLine(result.GetProperty("text").GetString());
foreach (var w in result.GetProperty("words").EnumerateArray())
    if (w.TryGetProperty("start", out var start))
        Console.WriteLine($"{start.GetDouble(),7:F2}s  {w.GetProperty("word").GetString()}");

What does the result contain?

From the STT 1.0 result schema in the Sume API reference and Sume's current worker code, read 2026-09-29.
FieldWhat it holds
textThe whole transcript
language_codeThe detected or requested language, when available
words[]word, plus start and end in seconds from the audio start when returned (current code); ordered by start; capped at 20,000 entries
segments[]Only with sentence segmentation: index, text, start, end, duration_seconds

What does it cost, and what can't it do?

STT 1.0 costs $0.01 per audio minute, plus a 5.5% agent fee by default, so a full 10-minute request is $0.10 before the fee.

  • No speaker labels: diarize is fixed server-side, and the current code refuses a request that sends it.
  • No microphone or live stream: the job reads a finished file. Real time speech to text API covers near-live chunking.
  • No SSE or WebSocket: poll, or pass a webhook_url for the terminal callback.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume