whisper-1 or gpt-transcribe for subtitles: what OpenAI assigns to each

OpenAI recommends gpt-transcribe, gpt-4o-transcribe-diarize for speakers, whisper-1 for translation and subtitles. Plus the 25 MB limit and a chunking script.

4 min readSume
All posts

OpenAI's speech to text guide recommends gpt-transcribe for general transcription, gpt-4o-transcribe-diarize when you need speaker labels, and whisper-1 for translation and subtitles. Uploads are limited to 25 MB per file.

For subtitles, that makes whisper-1 the model the guide points to, and the file limit makes chunking long audio a requirement.

Which model for which job

From the guide as read on 2026-10-03.

OpenAI transcription models by task (read 2026-10-03)
TaskModel the guide namesNotes
General transcriptiongpt-transcribeRecommended
Speaker labelsgpt-4o-transcribe-diarizediarized_json; up to 4 reference clips of 2 to 10 s
Translation and subtitleswhisper-1Named for these jobs
File sizeAll25 MB limit
StreamingListed optionstream=true

How long a file fits in 25 MB

The limit is on bytes, so duration depends on the encoding. Uncompressed 16-bit mono at 16 kHz is 32,000 bytes per second, so 25 MB (taking it as 25 million bytes) is about 781 seconds. That is arithmetic, not a figure from the guide. Compressed audio fits more minutes.

This script splits a file into fixed segments with ffmpeg, choosing a segment length that stays under the limit for the encoding you pass in.

import subprocess, sys

LIMIT = 25_000_000
BYTES_PER_SEC_WAV_16K_MONO = 16000 * 2

def segment_seconds(bytes_per_sec, margin=0.9):
    return int(LIMIT * margin / bytes_per_sec)

src = sys.argv[1]
seg = segment_seconds(BYTES_PER_SEC_WAV_16K_MONO)
subprocess.run([
    "ffmpeg", "-i", src, "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le",
    "-f", "segment", "-segment_time", str(seg), "part-%03d.wav",
], check=True)
print(f"segments of {seg} s written as part-NNN.wav")

Subtitle workflow notes

Choices that follow from the guide and the arithmetic.

  • Cutting on a fixed time can split a word. Keep a short overlap, or cut on silence, and merge the cues afterward.
  • Timestamps from each segment are relative to that segment. Add the segment's start offset before merging.
  • Run diarization separately if speaker labels matter, since the guide assigns that to a different model.
  • Keep the source audio and the model id with the subtitles so a corrected file can be regenerated.

Getting the audio out of a Sume video

If the video is a Sume-hosted artifact, POST /v1/audio-detach returns a new audio artifact. Sume's docs describe sample_rate: 16000 with channels: "mono" as the speech to text shape, at $0.01 per job, and note an output cap of 900 seconds, so longer tracks need a range.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume