whisper-1 or gpt-transcribe for subtitles: what OpenAI assigns to each
OpenAI recommends gpt-transcribe, gpt-4o-transcribe-diarize for speakers, whisper-1 for translation and subtitles. Plus the 25 MB limit and a chunking script.

OpenAI's speech to text guide recommends gpt-transcribe for general transcription, gpt-4o-transcribe-diarize when you need speaker labels, and whisper-1 for translation and subtitles. Uploads are limited to 25 MB per file.
For subtitles, that makes whisper-1 the model the guide points to, and the file limit makes chunking long audio a requirement.
Which model for which job
From the guide as read on 2026-10-03.
| Task | Model the guide names | Notes |
|---|---|---|
| General transcription | gpt-transcribe | Recommended |
| Speaker labels | gpt-4o-transcribe-diarize | diarized_json; up to 4 reference clips of 2 to 10 s |
| Translation and subtitles | whisper-1 | Named for these jobs |
| File size | All | 25 MB limit |
| Streaming | Listed option | stream=true |
How long a file fits in 25 MB
The limit is on bytes, so duration depends on the encoding. Uncompressed 16-bit mono at 16 kHz is 32,000 bytes per second, so 25 MB (taking it as 25 million bytes) is about 781 seconds. That is arithmetic, not a figure from the guide. Compressed audio fits more minutes.
This script splits a file into fixed segments with ffmpeg, choosing a segment length that stays under the limit for the encoding you pass in.
import subprocess, sys
LIMIT = 25_000_000
BYTES_PER_SEC_WAV_16K_MONO = 16000 * 2
def segment_seconds(bytes_per_sec, margin=0.9):
return int(LIMIT * margin / bytes_per_sec)
src = sys.argv[1]
seg = segment_seconds(BYTES_PER_SEC_WAV_16K_MONO)
subprocess.run([
"ffmpeg", "-i", src, "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le",
"-f", "segment", "-segment_time", str(seg), "part-%03d.wav",
], check=True)
print(f"segments of {seg} s written as part-NNN.wav")Subtitle workflow notes
Choices that follow from the guide and the arithmetic.
- Cutting on a fixed time can split a word. Keep a short overlap, or cut on silence, and merge the cues afterward.
- Timestamps from each segment are relative to that segment. Add the segment's start offset before merging.
- Run diarization separately if speaker labels matter, since the guide assigns that to a different model.
- Keep the source audio and the model id with the subtitles so a corrected file can be regenerated.
Getting the audio out of a Sume video
If the video is a Sume-hosted artifact, POST /v1/audio-detach returns a new audio artifact. Sume's docs describe sample_rate: 16000 with channels: "mono" as the speech to text shape, at $0.01 per job, and note an output cap of 900 seconds, so longer tracks need a range.
Sources
Related posts
More in Developers
- Windsurf now redirects to Devin Desktop: where Sume MCP setup lives
Windsurf redirects to Devin Desktop, and Cascade was removed in v3.9.19. Re-add Sume's hosted MCP URL there and verify it with mcp_health and tools_list.
- Which MCP server lets Claude Code or Cursor generate video and images?
MCP servers that let Claude Code and Cursor make video and images: Sume, fal, Replicate, Runway, Higgsfield. Endpoints, sign-in, billing, setup.
- Idempotency keys for AI video APIs: retry without paying twice
An idempotency key makes a retried create return the original run or job instead of a second paid one. How Sume's Idempotency-Key works on each API.
- Signed webhooks for Sume video runs: events, retries, verification
Sume sends one HMAC-SHA256 signed POST when a Format, Action, or Agent Completion run completes or fails. Verify the raw body and dedupe on request_id.
Written by Sume