Model family

OpenAIUpdated 2026AudioTranscriptionSpeech recognitionLocal

Whisper Models

Whisper models are used for ASR, transcription, subtitles, podcast processing, meeting notes, and multilingual audio.

Best for

Audio

Use this family hub to compare Whisper variants for audio workflows, then open the detail page for deeper deployment notes.

Transcription

Use this family hub to compare Whisper variants for transcription workflows, then open the detail page for deeper deployment notes.

Speech recognition

Use this family hub to compare Whisper variants for speech recognition workflows, then open the detail page for deeper deployment notes.

Local

Use this family hub to compare Whisper variants for local workflows, then open the detail page for deeper deployment notes.

Source box

This family currently includes 9 records tied to an exact published checkpoint. Identity is recorded per model so a representative checkpoint is never treated as the whole family.

Identity checked: 2026-08-21

Artifact identity does not establish licence or context truth. Those checks remain separate.

Variants

Whisper models grouped by workflow

Audio

AudioSpeech recognitionTranscriptionOpen source weights and code

Whisper Large V3

OpenAI · Whisper

Best for: Builders adding local transcription, podcast processing, meeting notes, or audio translation.

Local: Feed it 16 kHz mono audio. Transcription and speech-to-English translation are the same model with a different task token, so no second checkpoint is needed for translation.

Details →
AudioOpen weights where releasedaudiotranscription

Whisper Large V3 Turbo

OpenAI · Whisper

Best for: Fast transcription and multilingual audio workflows

Local: The usual default for interactive or real-time-ish transcription: pick it first, and only step up to large-v3 if measured accuracy on your own audio falls short.

Details →
AudioOpen source weights and codeaudiotranscription

Whisper Large V2

OpenAI · Whisper

Best for: Teams running accuracy-first transcription, subtitles, meeting notes, or archive workflows on multilingual audio.

Local: Swapping to large-v3 is a checkpoint change with no pipeline change, so an A/B on your own audio is cheap. Note that v3 expects 128-bin features, which the processor handles automatically.

Details →
AudioOpen source weights and codeaudiotranscription

Whisper Medium

OpenAI · Whisper

Best for: Builders balancing local transcription quality and runtime cost for meetings, media, and batch audio processing.

Local: Worth benchmarking against large-v3-turbo before committing: on similar hardware the turbo model usually wins on both speed and accuracy.

Details →
AudioOpen source weights and codeaudiotranscription

Whisper Small

OpenAI · Whisper

Best for: Local transcription setups that need a lighter model for captions, notes, and general speech-to-text tasks.

Local: A good fit for on-device or embedded transcription where the alternative is sending audio off the machine.

Details →
AudioOpen source weights and codeaudiotranscription

Whisper Base

OpenAI · Whisper

Best for: Developers who want a practical starting point for local captioning, voice notes, and CPU-leaning ASR experiments.

Local: A sensible first checkpoint when wiring up an audio pipeline: get the plumbing working here, then swap in a larger size without changing any code.

Details →
AudioOpen source weights and codeaudiotranscription

Whisper Tiny

OpenAI · Whisper

Best for: Fast prototypes, edge-style experiments, and low-memory ASR tests where accuracy tradeoffs are acceptable.

Local: Useful for voice-activity and language-detection steps in front of a larger model, where being wrong about words costs nothing.

Details →
AudioOpen weightsaudiotranscription

Distil-Whisper Large V3

Hugging Face / community · Whisper

Best for: Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.

Local: The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.

Details →
AudioOpen weights where releasedaudiotranscription

Faster-Whisper Large V3

SYSTRAN / community · Whisper

Best for: Teams optimizing Whisper-style batch or local transcription pipelines where runtime efficiency and deployment control matter.

Local: The pragmatic way to run large-v3 in production: same accuracy, materially better throughput, with word-level timestamps and voice-activity filtering handled by the runtime.

Details →

Compare

All Whisper models in the directory

ModelTypeBest forLocal runner notesLicenseDetail
Whisper Large V3AudioBuilders adding local transcription, podcast processing, meeting notes, or audio translation.Feed it 16 kHz mono audio. Transcription and speech-to-English translation are the same model with a different task token, so no second checkpoint is needed for translation.Apache 2.0Open
Whisper Large V3 TurboAudioFast transcription and multilingual audio workflowsThe usual default for interactive or real-time-ish transcription: pick it first, and only step up to large-v3 if measured accuracy on your own audio falls short.MITOpen
Whisper Large V2AudioTeams running accuracy-first transcription, subtitles, meeting notes, or archive workflows on multilingual audio.Swapping to large-v3 is a checkpoint change with no pipeline change, so an A/B on your own audio is cheap. Note that v3 expects 128-bin features, which the processor handles automatically.Apache 2.0Open
Whisper MediumAudioBuilders balancing local transcription quality and runtime cost for meetings, media, and batch audio processing.Worth benchmarking against large-v3-turbo before committing: on similar hardware the turbo model usually wins on both speed and accuracy.Apache 2.0Open
Whisper SmallAudioLocal transcription setups that need a lighter model for captions, notes, and general speech-to-text tasks.A good fit for on-device or embedded transcription where the alternative is sending audio off the machine.Apache 2.0Open
Whisper BaseAudioDevelopers who want a practical starting point for local captioning, voice notes, and CPU-leaning ASR experiments.A sensible first checkpoint when wiring up an audio pipeline: get the plumbing working here, then swap in a larger size without changing any code.Apache 2.0Open
Whisper TinyAudioFast prototypes, edge-style experiments, and low-memory ASR tests where accuracy tradeoffs are acceptable.Useful for voice-activity and language-detection steps in front of a larger model, where being wrong about words costs nothing.Apache 2.0Open
Distil-Whisper Large V3AudioBuilders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.MITOpen
Faster-Whisper Large V3AudioTeams optimizing Whisper-style batch or local transcription pipelines where runtime efficiency and deployment control matter.The pragmatic way to run large-v3 in production: same accuracy, materially better throughput, with word-level timestamps and voice-activity filtering handled by the runtime.Apache 2.0 (weights); MIT (faster-whisper code)Open

Source box

This family currently includes 9 records tied to an exact published checkpoint. Identity is recorded per model so a representative checkpoint is never treated as the whole family.

Identity checked: 2026-08-21

Artifact identity does not establish licence or context truth. Those checks remain separate.