Keyterm prompting costs $0.05/hour at two vendors: 60 hours is $3.00
ElevenLabs and AssemblyAI both list keyterm prompting at $0.05 an hour: $3.00 for 60 hours of calls. Sume STT has a language hint but no keyterm list.

ElevenLabs lists Scribe keyterm prompting at $0.050 per hour and AssemblyAI lists keyterms prompting for asynchronous audio at +$0.05 per hour. For 60 hours of product calls, either add-on costs 60 x $0.05 = $3.00 on top of the base transcription. Sume STT takes a language hint but no keyterm list, so there is no equivalent line, and it is priced at $0.01 a minute, or $36.00 for 60 hours.
Rates were read on each vendor's page and on Sume's catalog on 2026-10-09.
What keyterms are for
A keyterm list tells the model which uncommon words to expect, such as product names, people's names or jargon, so it spells them correctly instead of guessing a common word that sounds alike. For a support team that talks about a product with an invented name, the gain is fewer corrections afterward.
| Option | Per hour | 60 hours |
|---|---|---|
| ElevenLabs Scribe v2 | $0.22 | $13.20 |
| ElevenLabs Scribe v2 plus keyterm prompting | $0.27 | $16.20 |
| AssemblyAI Universal-3.5 Pro | $0.21 | $12.60 |
| AssemblyAI Universal-3.5 Pro plus keyterms | $0.26 | $15.60 |
| Sume STT 1.0 (no keyterm option) | $0.60 | $36.00 |
What to do with Sume STT
Sume's STT request has audio_url, language_code, duration_seconds, a segmentation block and the usual job fields. Provider knobs are fixed server-side, so you cannot send a keyterm list. If the proper nouns matter, correct them after transcription with a find-and-replace table built from the common misspellings you see in a sample.
When the transcript feeds captions and you already have the intended script, the video caption job offers script_text. Sume keeps the speech-to-text word timings and aligns the burned-in text to your script, so product names appear as you wrote them. That route needs the script, which a support call does not have.
Is $3.00 worth it
On 60 hours, $3.00 is trivial next to the cost of reading the transcripts. Test on one hour first: transcribe it with and without keyterms and count the corrections. If the add-on saves more than a couple of minutes of editing per hour, it has paid for itself at any normal wage.
The more useful comparison is base rate. On these pages Sume's price per hour is the highest of the four rows, which is why this post is about what you give up, not about a saving.
- Vendors with keyterms: $0.05 per hour on top of the base price.
- Sume: no keyterm field; fix names after the fact or align to a script.
- Always test with one real hour before committing 60.
Building a keyterm list
A keyterm list works best when it is short and specific. Collect the 20 or 30 words your team gets wrong most often: product names, feature names, competitor names, and any customer names you are allowed to store. Do not add common words, since that adds noise without helping.
Review the list monthly. A new product launch changes the vocabulary of every call, and a stale list can bias the model toward a name that no longer comes up. Keep the list in version control next to your transcription script so a change is visible to everyone who relies on the output.
Sources
Related posts
More in Comparisons
- Kling 3 Pro text-to-video is $0.14/s on fal; four Sume ids cost less
Sume lists no Kling text-to-video, only Kling 3.0 Motion Control at $0.1575. Four Sume text-to-video ids cost $0.075 to $0.125 per second against fal's $0.14.
- LTX-2.5 vs MiniMax H3: license lines and run requirements
LTX-2.5 vs MiniMax H3 from vendor pages: 10M vs 20M USD revenue lines, excluded territories, Python and CUDA needs, frame rules, audio, and what Sume lists.
- Luma Ray 3.2 alternative: one /v1/generations route vs Sume's two
Luma's API sends images and video through POST /v1/generations. Sume splits them into /v1/videos and /v1/images and does not list ray-3.2 in its docs.
- Luma Ray3.2 allows 16 keyframes; Sume offers first and last frame
Luma lists up to 16 keyframes per clip on its Ray3.2 API. Sume does not list Luma; its video models take a first and a last frame through frame_images.
Written by Sume