- Models
- Whisper
Model family
Whisper Models
Whisper models are used for ASR, transcription, subtitles, podcast processing, meeting notes, and multilingual audio.
Best for
Audio
Use this family hub to compare Whisper variants for audio workflows, then open the detail page for deeper deployment notes.
Transcription
Use this family hub to compare Whisper variants for transcription workflows, then open the detail page for deeper deployment notes.
Speech recognition
Use this family hub to compare Whisper variants for speech recognition workflows, then open the detail page for deeper deployment notes.
Local
Use this family hub to compare Whisper variants for local workflows, then open the detail page for deeper deployment notes.
Source box
This family currently includes 9 records tied to an exact published checkpoint. Identity is recorded per model so a representative checkpoint is never treated as the whole family.
Identity checked: 2026-08-21
Artifact identity does not establish licence or context truth. Those checks remain separate.
Jump to
Variants
Whisper models grouped by workflow
Audio
Whisper Large V3
OpenAI · Whisper
Best for: Builders adding local transcription, podcast processing, meeting notes, or audio translation.
Local: Feed it 16 kHz mono audio. Transcription and speech-to-English translation are the same model with a different task token, so no second checkpoint is needed for translation.
Whisper Large V3 Turbo
OpenAI · Whisper
Best for: Fast transcription and multilingual audio workflows
Local: The usual default for interactive or real-time-ish transcription: pick it first, and only step up to large-v3 if measured accuracy on your own audio falls short.
Whisper Large V2
OpenAI · Whisper
Best for: Teams running accuracy-first transcription, subtitles, meeting notes, or archive workflows on multilingual audio.
Local: Swapping to large-v3 is a checkpoint change with no pipeline change, so an A/B on your own audio is cheap. Note that v3 expects 128-bin features, which the processor handles automatically.
Whisper Medium
OpenAI · Whisper
Best for: Builders balancing local transcription quality and runtime cost for meetings, media, and batch audio processing.
Local: Worth benchmarking against large-v3-turbo before committing: on similar hardware the turbo model usually wins on both speed and accuracy.
Whisper Small
OpenAI · Whisper
Best for: Local transcription setups that need a lighter model for captions, notes, and general speech-to-text tasks.
Local: A good fit for on-device or embedded transcription where the alternative is sending audio off the machine.
Whisper Base
OpenAI · Whisper
Best for: Developers who want a practical starting point for local captioning, voice notes, and CPU-leaning ASR experiments.
Local: A sensible first checkpoint when wiring up an audio pipeline: get the plumbing working here, then swap in a larger size without changing any code.
Whisper Tiny
OpenAI · Whisper
Best for: Fast prototypes, edge-style experiments, and low-memory ASR tests where accuracy tradeoffs are acceptable.
Local: Useful for voice-activity and language-detection steps in front of a larger model, where being wrong about words costs nothing.
Distil-Whisper Large V3
Hugging Face / community · Whisper
Best for: Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
Local: The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.
Faster-Whisper Large V3
SYSTRAN / community · Whisper
Best for: Teams optimizing Whisper-style batch or local transcription pipelines where runtime efficiency and deployment control matter.
Local: The pragmatic way to run large-v3 in production: same accuracy, materially better throughput, with word-level timestamps and voice-activity filtering handled by the runtime.
Compare
All Whisper models in the directory
| Model | Type | Best for | Local runner notes | License | Detail |
|---|---|---|---|---|---|
| Whisper Large V3 | Audio | Builders adding local transcription, podcast processing, meeting notes, or audio translation. | Feed it 16 kHz mono audio. Transcription and speech-to-English translation are the same model with a different task token, so no second checkpoint is needed for translation. | Apache 2.0 | Open |
| Whisper Large V3 Turbo | Audio | Fast transcription and multilingual audio workflows | The usual default for interactive or real-time-ish transcription: pick it first, and only step up to large-v3 if measured accuracy on your own audio falls short. | MIT | Open |
| Whisper Large V2 | Audio | Teams running accuracy-first transcription, subtitles, meeting notes, or archive workflows on multilingual audio. | Swapping to large-v3 is a checkpoint change with no pipeline change, so an A/B on your own audio is cheap. Note that v3 expects 128-bin features, which the processor handles automatically. | Apache 2.0 | Open |
| Whisper Medium | Audio | Builders balancing local transcription quality and runtime cost for meetings, media, and batch audio processing. | Worth benchmarking against large-v3-turbo before committing: on similar hardware the turbo model usually wins on both speed and accuracy. | Apache 2.0 | Open |
| Whisper Small | Audio | Local transcription setups that need a lighter model for captions, notes, and general speech-to-text tasks. | A good fit for on-device or embedded transcription where the alternative is sending audio off the machine. | Apache 2.0 | Open |
| Whisper Base | Audio | Developers who want a practical starting point for local captioning, voice notes, and CPU-leaning ASR experiments. | A sensible first checkpoint when wiring up an audio pipeline: get the plumbing working here, then swap in a larger size without changing any code. | Apache 2.0 | Open |
| Whisper Tiny | Audio | Fast prototypes, edge-style experiments, and low-memory ASR tests where accuracy tradeoffs are acceptable. | Useful for voice-activity and language-detection steps in front of a larger model, where being wrong about words costs nothing. | Apache 2.0 | Open |
| Distil-Whisper Large V3 | Audio | Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines. | The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint. | MIT | Open |
| Faster-Whisper Large V3 | Audio | Teams optimizing Whisper-style batch or local transcription pipelines where runtime efficiency and deployment control matter. | The pragmatic way to run large-v3 in production: same accuracy, materially better throughput, with word-level timestamps and voice-activity filtering handled by the runtime. | Apache 2.0 (weights); MIT (faster-whisper code) | Open |