Hypit transcribe cost: $0.1082 CPU hold vs $0.1758 on a T4
Hypit transcribe holds $0.1082 for WhisperX small on CPU and $0.1758 for medium or large-v3 on a T4 GPU. Scribe v2 makes no Modal call. How the holds are built.

Hypit transcribe holds **$0.1082** when it runs WhisperX small on CPU and **$0.1758** when it runs medium or large-v3 on a Modal T4. Both are ceilings over 600 seconds, priced at Modal list x 1.25 plus the 5.5% fee. engine: scribe_v2 makes no Modal call and keeps its own speech-to-text hold.
The two shapes
A GPU container also pays for its cores and memory. Modal's page lists a T4 at $0.000164 per second, on top of $0.0000131 per core-second and $0.00000222 per GiB-second.
| Ceiling | Model | Shape x seconds | Hold |
|---|---|---|---|
| hypit_transcribe_cpu | WhisperX small | 8 cores / 8 GiB + 1 core / 512 MiB control, x 600 s | $0.1082 |
| hypit_transcribe_gpu | WhisperX medium, large-v3* | T4 + 2 cores / 8 GiB + 1 core / 512 MiB control, x 600 s | $0.1758 |
Why there is a control container
Each transcribe call runs a small control container next to the worker. It is billed too, which is why the shape lists two lines. The control container is tiny, 1 core and 512 MiB, so it adds well under a cent.
Which one to pick
Choose the CPU path for short clips, drafts and clean audio. Move to the T4 shapes when accuracy on noisy or accented speech matters more than the extra cents. The difference between the two holds is $0.0676 per call at the ceiling.
If you need a cloud transcript with no Modal line at all, engine: scribe_v2 uses the Scribe v2 route at the published sume/stt-1.0 price of $0.01 per audio minute.
What is capped
The charge never exceeds the hold. A clipped capture is flagged price_book_clipped_to_hold: 1 on the usage row, so a surprising bill can be traced to a ceiling rather than a pricing change.
How the T4 hold is assembled
Modal lists a T4 at $0.000164 per second. Over 600 seconds that is $0.0984 before any cores or memory. Add the 2 cores and 8 GiB of the worker and the 1 core and 512 MiB of the control container, then x 1.25 and the 5.5% fee, and you arrive at the $0.1758 hold.
The CPU path has no GPU line, which is why its hold is lower even though its worker has 8 cores and 8 GiB.
Choosing between WhisperX and Scribe v2
The transcript quality tradeoffs, such as word confidence, are covered in the dedicated WhisperX versus Scribe v2 comparison. For cost, the rule is simple. WhisperX is a Modal line that scales with the container seconds used. Scribe v2 is a per-minute line at $0.01 per audio minute and avoids the container entirely.
Decide on accuracy and features first, then check the hold.
Sources
Related posts
More in Media tools
- Link card 1200x628 from an Ideogram 4.5 edit: ask 2:1, crop 63 px
Ideogram 4.5 on Sume has no 1.91:1 ratio. Ask for 2:1 (1408x704 on a 1K edit), then crop about 63 px of width with Pillow to reach 1200x628.
- ImageMagick: turn a Sume WebP into a 1080x1350 JPEG
One magick command resizes to fill, center-crops to 1080x1350 and writes a stripped JPEG. Use it on a Sume 4:5 image whose native size is not exactly 1080x1350.
- Trim a TikTok or Instagram clip on Sume: import first, then cut
Sume media tools only read files on media.sume.com. To trim a public TikTok or Instagram video, import it first, then send the artifact URL to video-trim.
- Media import allows about 2 GB, but trim and filter stop at 300 MiB
Sume media import has a practical size cap near 2 GB, while video trim and filter reject sources over 300 MiB with source_too_large. How to plan around the gap.
Written by Sume