IBM Watson Speech to Text Lite: 500 free minutes vs Sume STT $5
IBM's Lite plan gives 500 free recognition minutes a month; Plus and Premium prices are not on the page. The same 500 minutes on Sume STT cost $5.00.

IBM's Speech to Text Lite plan includes 500 minutes of free speech recognition a month with 38 pre-trained models. The same 500 minutes on Sume STT cost $5.00 at $0.01 a minute. IBM's pricing page does not show per-minute prices for the paid Plus and Premium plans, so a price comparison above the free tier needs IBM's detailed pricing documentation or its sales team.
IBM's tiers are from its pricing page, read 2026-10-03.
What IBM lists
The page names three plans and gives concrete numbers only for the first.
| Plan | Included minutes | Concurrency | Other |
|---|---|---|---|
| Lite | 500 free a month | Not stated | 38 pre-trained speech models |
| Plus | Unlimited | 100 concurrent transcriptions | Model tuning |
| Premium | Unlimited | Unlimited | For large and security-sensitive organizations |
Free tier math
Five hundred minutes is 8 hours 20 minutes. On Sume that is $5.00, spread over 50 jobs of 10 minutes at $0.10 each. If your volume is a few meetings a month, IBM's free tier covers it and costs nothing. Once you exceed it, the page tells you to check IBM's own price documentation, and I did not read that.
Plus and Premium describe unlimited monthly minutes, which suggests a fee structure that is not per-minute, but the page I read does not give the fee, so I will not guess.
Volumes against the free tier
Sume has no free tier in the docs I read, so every minute is billed. IBM Lite is free up to 500 and the page does not price what comes after on the paid plans.
| Minutes a month | IBM Lite | Sume STT |
|---|---|---|
| 300 | $0 | $3.00 |
| 500 | $0 | $5.00 |
| 1,000 | Needs a paid plan, price not shown | $10.00 |
What Sume does and does not do
Sume STT takes a public HTTPS audio URL and returns a transcript; omit the language code for auto-detect. The cap is 600 seconds per job. It has no model tuning and no custom vocabulary on the request schema I read. See the API reference.
Speech to text is $0.01 per minute of audio, up to 600 seconds per job.
- If you need custom models or very high concurrency, IBM's Plus tier is the one built for that.
- If you need transcripts feeding captions, clips or dubbing in the same API, Sume keeps them together.
- For a long file see splitting recordings.
Sources
Related posts
More in Comparisons
- Ideogram 4.5 Magic Fill and Extend vs Sume mask_url and aspect ratio
Ideogram's docs list Magic Fill and Extend as editing features. Sume has no Extend endpoint; here is what its mask and aspect-ratio fields cover instead.
- Ideogram 4.5 vs Ideogram V3 on the Sume image API
Ideogram V3 lists $0.075 with up to 10 references; Ideogram 4.5 lists $0.075 at medium quality with 5 references and a 1K or 2K resolution field.
- Instagram's Devanagari caption fonts: what Sume burns for Hindi
Instagram added Devanagari and Bengali-Assamese caption fonts. Sume documents Latin and Hangul faces only, so test Hindi cues on one clip first.
- Reels run to 20 minutes; Sume's audio and caption limits to plan for
Instagram says Reels can reach 20 minutes but over 3 minutes is not recommended. The Sume Timeline audio, avatar and caption limits that matter at that length.
Written by Sume