Audio
Distil-Whisper Large V3
Distilled Whisper large-v3 checkpoint for Whisper-style transcription workflows with a smaller deployment footprint than the full large model.
Hugging Face / community · Whisper
Editorial review
Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.
Best for
Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
Who should use it
- Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
- Builders who want local or self-hosted testing options.
Common workflows
- Distilled transcription workflows
- audio workflows
- transcription workflows
- speech recognition workflows
- local workflows
Deployment and hardware notes
~1.5 GB in fp16, comparable to whisper-medium, but decoding is far cheaper because only two decoder layers run per token.
License and usage notes
MIT. Open weights. Verify the exact model card and license terms for the checkpoint or hosted provider you use.
Strengths
- Open weights model option for Whisper workflows.
- Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
- The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.
Limitations
- English only, which is the trade that buys the speed: it cannot transcribe or translate other languages at all, where every OpenAI Whisper checkpoint covers 99. Its licence is MIT rather than Apache 2.0.
- ~1.5 GB in fp16, comparable to whisper-medium, but decoding is far cheaper because only two decoder layers run per token.
- Context window and limits: 30-second audio windows · 448-token cap per window.
- Verify the exact model card, provider docs, license, and serving support before production use.
Local workflow notes
The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.
Local runtimes: Transformers, faster-whisper, whisper.cpp
Platforms: Windows, macOS, Linux
Transcription spec
A distillation of large-v3 down to a 2-layer decoder — shallower than large-v3-turbo's 4 — with the 32-layer encoder kept intact. The fastest decoder on this page by a clear margin.
Sources to verify
Related resources
Continue with model source notes, local tools, and implementation guides related to this model.
Model ecosystem connections
Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.
Recommended runtimes and tools
Setup and deployment
Related model pages
Guides, stacks, and comparisons
Ready to run this model locally?
Find a compatible interface in our Local AI Tools directory →