Hypit transcribe cost: $0.1082 CPU hold vs $0.1758 on a T4

Hypit transcribe holds $0.1082 for WhisperX small on CPU and $0.1758 for medium or large-v3 on a T4 GPU. Scribe v2 makes no Modal call. How the holds are built.

5 min readSume
All posts

Hypit transcribe holds **$0.1082** when it runs WhisperX small on CPU and **$0.1758** when it runs medium or large-v3 on a Modal T4. Both are ceilings over 600 seconds, priced at Modal list x 1.25 plus the 5.5% fee. engine: scribe_v2 makes no Modal call and keeps its own speech-to-text hold.

The two shapes

A GPU container also pays for its cores and memory. Modal's page lists a T4 at $0.000164 per second, on top of $0.0000131 per core-second and $0.00000222 per GiB-second.

Holds from the Sume price book; Modal rates from modal.com/pricing, read 2026-10-05.
CeilingModelShape x secondsHold
hypit_transcribe_cpuWhisperX small8 cores / 8 GiB + 1 core / 512 MiB control, x 600 s$0.1082
hypit_transcribe_gpuWhisperX medium, large-v3*T4 + 2 cores / 8 GiB + 1 core / 512 MiB control, x 600 s$0.1758

Why there is a control container

Each transcribe call runs a small control container next to the worker. It is billed too, which is why the shape lists two lines. The control container is tiny, 1 core and 512 MiB, so it adds well under a cent.

Which one to pick

Choose the CPU path for short clips, drafts and clean audio. Move to the T4 shapes when accuracy on noisy or accented speech matters more than the extra cents. The difference between the two holds is $0.0676 per call at the ceiling.

If you need a cloud transcript with no Modal line at all, engine: scribe_v2 uses the Scribe v2 route at the published sume/stt-1.0 price of $0.01 per audio minute.

What is capped

The charge never exceeds the hold. A clipped capture is flagged price_book_clipped_to_hold: 1 on the usage row, so a surprising bill can be traced to a ceiling rather than a pricing change.

How the T4 hold is assembled

Modal lists a T4 at $0.000164 per second. Over 600 seconds that is $0.0984 before any cores or memory. Add the 2 cores and 8 GiB of the worker and the 1 core and 512 MiB of the control container, then x 1.25 and the 5.5% fee, and you arrive at the $0.1758 hold.

The CPU path has no GPU line, which is why its hold is lower even though its worker has 8 cores and 8 GiB.

Choosing between WhisperX and Scribe v2

The transcript quality tradeoffs, such as word confidence, are covered in the dedicated WhisperX versus Scribe v2 comparison. For cost, the rule is simple. WhisperX is a Modal line that scales with the container seconds used. Scribe v2 is a per-minute line at $0.01 per audio minute and avoids the container entirely.

Decide on accuracy and features first, then check the hold.

Sources

Related posts

More in Media tools

All Media tools posts

Written by Sume