NVIDIA Parakeet v3 or Canary v2 vs a hosted STT API: Sume STT 1.0
Parakeet TDT v3 and Canary 1B v2 are open ASR models for 25 European languages. When to self-host them and when Sume STT at $0.01 a minute is enough.

Self-host NVIDIA Parakeet TDT 0.6B v3 or Canary 1B v2 if you need open weights, your own GPUs and a European-language model you can modify; call Sume STT 1.0 if you want a hosted transcript with word timings at $0.01 per audio minute and no GPU to run. The two are not rivals so much as different ends of the same buy-or-run decision.
Both NVIDIA cards were read on 2026-10-10. Nothing below compares accuracy between them and Sume, because Sume publishes no benchmark and none is invented here.
What the two NVIDIA model cards say
Parakeet TDT 0.6B v3 is a 600-million-parameter model for 25 European languages under CC-BY-4.0. The card lists automatic punctuation and word and segment timestamps, a 16 kHz mono wav or flac input, and a reported average WER of 7.83% on MLS, 1.93% on LibriSpeech clean and 6.34% on the Hugging Face Open ASR Leaderboard average. It handles up to 24 minutes with full attention on an A100 80GB, or up to three hours with local attention.
Canary 1B v2 covers the same 25 languages and adds translation: any of the languages into English, and English into the other 24. Its card reports 8.40% WER on Fleurs-25, 8.85% on CoVoST-13 and 7.27% on MLS-6, also under CC-BY-4.0 with 16 kHz mono input. Word timestamps come with transcription; translation returns segment timestamps only.
| Item | Parakeet TDT 0.6B v3 | Canary 1B v2 |
|---|---|---|
| Languages | 25 European | 25 European |
| Task | Transcription | Transcription plus translation |
| License | CC-BY-4.0 | CC-BY-4.0 |
| Input | 16 kHz mono wav or flac | 16 kHz mono wav or flac |
| Timestamps | Word and segment | Word (transcription), segment (translation) |
| Both released | Aug 14, 2025 | Aug 14, 2025 |
What Sume STT 1.0 does instead
Sume STT 1.0 is one POST: send a public HTTPS audio_url, optionally a language_code (omit it to auto-detect), and read back text plus word-level timings in seconds. Add segmentation with mode sentence to get sentence groups derived from those word timings. Billing is $0.01 per audio minute, and an optional duration_seconds hint of 1 to 600 helps usage reservation; omit it and the job reserves one minute.
The 600-second ceiling on that hint means a long recording is submitted in chunks of up to ten minutes. The default mode is async; sync waits at most 30 seconds, so use async or a webhook for anything real.
Making a Sume video fit an open model's input
Both NVIDIA cards want 16 kHz mono audio. If your source is a video on Sume, the audio detach job extracts the track as wav with channels set to mono and sample_rate 16000, for $0.01 per job. That is the same shape used for STT, so one extraction can feed Sume STT or a model you run yourself.
Audio detach takes a source of at most 1,800 seconds and writes at most 900 seconds of output, and the video must already be on media.sume.com. Longer files need ranges, as in the three-range talk example.
How to choose
The deciding questions are about control and volume, not a leaderboard number.
- Choose the open models when audio cannot leave your network, when you need to fine-tune, or when you already own idle GPU capacity.
- Choose Canary when you need translation to English in the same model; Sume STT returns a transcript in the spoken language and does not translate.
- Choose Sume STT when you want no infrastructure, a signed job envelope, webhooks, and a flat per-minute price: one hour of audio is 60 minutes, so $0.60 at the list rate before any rounding.
- Check language coverage first. The NVIDIA cards list 25 European languages; Sume lets you pass a language hint and auto-detects otherwise, but it does not publish a language list in the sources read for this post.
What this post cannot tell you
It does not compare accuracy, latency or total cost of ownership. GPU hours, engineering time and queueing are yours to price for the self-hosted route, and the NVIDIA WER figures come from NVIDIA's own test sets, not from your audio.
A fair test is to run the same ten recordings through both and compare the words you care about. For the hosted side, the extraction recipe is in audio detach for speech-to-text; a language hint example is in Arabic speech-to-text on Sume.
Sources
Related posts
More in Comparisons
- Poster headline in the image: Ideogram 4.5 or Nano Banana 2.1 on Sume?
Both vendors pitch readable text in images. On Sume, Ideogram 4.5 starts at $0.0375 and Nano Banana 2.1 at $0.075. Compare price, ratios and reference limits.
- Qwen Image Max vs Qwen Image on Sume: 3.75x the price, no references
On Sume, Qwen Image Max bills $0.09375 per image and is text-only, while Qwen Image bills $0.025 and takes up to 10 references. Both list 13 aspect ratios.
- Recraft V4 vs FLUX.2 Pro on Sume: $0.05 text-only vs $0.0375 with refs
On Sume, Recraft V4 bills $0.05 per image, is text-only and returns webp only. FLUX.2 Pro bills $0.0375 and takes 10 references. Both list 13 ratios.
- Remotion Automator $0.01 a render vs Sume Timeline $0.10 a minute
Remotion lists $0.01 a render with a $100 monthly minimum; Sume Timeline bills $0.10 per output minute. Break-even by volume and what each leaves out.
Written by Sume