Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01

Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.

5 min readSume
All posts

Together AI's pricing page lists Whisper Large v3 transcription at $0.0015 per audio minute. Sume's STT 1.0 is $0.01 per audio minute, about 6.7 times as much, with a 10-minute maximum per request. On 1,000 minutes that is $1.50 against $10.00.

The two are not the same service, so the gap is a number to weigh, not a verdict. This post gives the rows and the arithmetic, and says where Sume does not match what Together lists.

What does Together AI list for transcription?

The speech-to-text table on Together AI's pricing page, read on 2026-10-03, prices per audio minute and notes that a Batch API price is available.

Transcription price per audio minute (read 2026-10-03)
ServicePer minute1,000 minutes10,000 minutes
Together: Whisper Large v3$0.0015$1.50$15.00
Together: NVIDIA Parakeet TDT 0.6B v3$0.0015$1.50$15.00
Together: Whisper Large v3 (Streaming)$0.0035$3.50$35.00
Together: NVIDIA Nemotron 3.5 ASR$0.0045$4.50$45.00
Sume STT 1.0$0.01$10.00$100.00

How does Sume bill transcription?

Sume's catalog prices STT 1.0 per audio minute at $0.01 after margin. A request reserves from duration_seconds when you send it, or one minute when you omit it, and the maximum is 10 minutes per request. The call is POST /v1/stt-1.0/transcribe, listed in the timeline audio docs.

Reserve, capture and refund follow the generation admission rules: the hold is placed at submit and settled when the job finishes. A failed transcription is refunded.

What does the cap mean for long audio?

A one-hour recording is six requests at the 10-minute maximum, so you split the audio yourself and join the text. That is $0.60 for the hour on Sume against $0.09 on Together's Whisper row. The Sume docs I read do not describe a long-audio mode, so if your input is meetings or podcasts, count the chunking work as part of the cost.

Sume STT also does not label speakers; the Apple Podcasts transcript post covers that limit. I did not check whether the Together models label speakers, because the pricing page does not say.

When does the gap matter?

At 1,000 minutes the difference is $8.50. For a team transcribing the hook line of ad clips, the transcript is a rounding error next to the picture: a 15-second Avatar Plus clip is $3.675 and its transcript is $0.01. At call-centre scale of 100,000 minutes a month the gap is $850, and a dedicated transcription vendor is the sensible choice.

Sume's reason to transcribe is the pipeline: the transcript can feed captions or the timeline in the same wallet, with one balance to watch. If transcription is all you need, price it where it is cheapest.

What would change the comparison?

Different audio length, a batch price on Together's side, or accuracy needs that favour one model. This post tests none of those. Re-read both pages before committing volume: they are dated 2026-10-03 here.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume