Together AI speech-to-text at $0.0015 a minute vs Sume STT at $0.01
Together AI lists Whisper Large v3 at $0.0015 per audio minute. Sume STT is $0.01 per minute with a 10-minute cap. Cost of 1,000 minutes, and what the gap buys.

Together AI's pricing page lists Whisper Large v3 transcription at $0.0015 per audio minute. Sume's STT 1.0 is $0.01 per audio minute, about 6.7 times as much, with a 10-minute maximum per request. On 1,000 minutes that is $1.50 against $10.00.
The two are not the same service, so the gap is a number to weigh, not a verdict. This post gives the rows and the arithmetic, and says where Sume does not match what Together lists.
What does Together AI list for transcription?
The speech-to-text table on Together AI's pricing page, read on 2026-10-03, prices per audio minute and notes that a Batch API price is available.
| Service | Per minute | 1,000 minutes | 10,000 minutes |
|---|---|---|---|
| Together: Whisper Large v3 | $0.0015 | $1.50 | $15.00 |
| Together: NVIDIA Parakeet TDT 0.6B v3 | $0.0015 | $1.50 | $15.00 |
| Together: Whisper Large v3 (Streaming) | $0.0035 | $3.50 | $35.00 |
| Together: NVIDIA Nemotron 3.5 ASR | $0.0045 | $4.50 | $45.00 |
| Sume STT 1.0 | $0.01 | $10.00 | $100.00 |
How does Sume bill transcription?
Sume's catalog prices STT 1.0 per audio minute at $0.01 after margin. A request reserves from duration_seconds when you send it, or one minute when you omit it, and the maximum is 10 minutes per request. The call is POST /v1/stt-1.0/transcribe, listed in the timeline audio docs.
Reserve, capture and refund follow the generation admission rules: the hold is placed at submit and settled when the job finishes. A failed transcription is refunded.
What does the cap mean for long audio?
A one-hour recording is six requests at the 10-minute maximum, so you split the audio yourself and join the text. That is $0.60 for the hour on Sume against $0.09 on Together's Whisper row. The Sume docs I read do not describe a long-audio mode, so if your input is meetings or podcasts, count the chunking work as part of the cost.
Sume STT also does not label speakers; the Apple Podcasts transcript post covers that limit. I did not check whether the Together models label speakers, because the pricing page does not say.
When does the gap matter?
At 1,000 minutes the difference is $8.50. For a team transcribing the hook line of ad clips, the transcript is a rounding error next to the picture: a 15-second Avatar Plus clip is $3.675 and its transcript is $0.01. At call-centre scale of 100,000 minutes a month the gap is $850, and a dedicated transcription vendor is the sensible choice.
Sume's reason to transcribe is the pipeline: the transcript can feed captions or the timeline in the same wallet, with one balance to watch. If transcription is all you need, price it where it is cheapest.
What would change the comparison?
Different audio length, a batch price on Together's side, or accuracy needs that favour one model. This post tests none of those. Re-read both pages before committing volume: they are dated 2026-10-03 here.
Sources
Related posts
More in Comparisons
- Together AI TTS runs $4 to $65 per million characters; Sume is $47.50
Together AI lists text-to-speech from $4 to $65 per million characters. Sume TTS is $0.0475 per 1,000, or $47.50 per million. Cost of 100 scripts on each.
- Veo 3.1 Lite vs Wan 3.0: price per second at 480p, 720p and 1080p
Google lists Veo 3.1 Lite at $0.05 (720p), $0.08 (1080p) per second; fal lists Wan 3.0 at $0.05 (480p), $0.10 (720p), $0.20 (1080p).
- Veo 3.1 tiers vs Seedance 2.5: price per second of video
Google lists Veo 3.1 at $0.40 (Standard), $0.10-$0.12 (Fast) and $0.05-$0.08 (Lite) per second; Seedance 2.5 on fal is about $0.46 at 720p. Sume x 1.25: $0.58.
- Wan 3.0 or MiniMax H3 Max: which to pin for reference-to-video
Both take image, video and audio references on Sume. Wan runs 2 to 30 seconds with 5 reference videos; H3 Max runs 5 to 15 seconds with always-on stereo audio.
Written by Sume