Audio

MITOpen weightsUpdated August 2026

Distil-Whisper Large V3

Distilled Whisper large-v3 checkpoint for Whisper-style transcription workflows with a smaller deployment footprint than the full large model.

Hugging Face / community · Whisper

Editorial review

Reviewed byOpenSourcesAI EditorialLast updatedAugust 2026SourcesExact model card

Model checkpoints, context windows, provider support, local runtime compatibility, and license terms can change quickly. Verify the exact model card before production or commercial use.

Best for

Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.

Who should use it

  • Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
  • Builders who want local or self-hosted testing options.

Common workflows

  • Distilled transcription workflows
  • audio workflows
  • transcription workflows
  • speech recognition workflows
  • local workflows

Deployment and hardware notes

~1.5 GB in fp16, comparable to whisper-medium, but decoding is far cheaper because only two decoder layers run per token.

License and usage notes

MIT. Open weights. Verify the exact model card and license terms for the checkpoint or hosted provider you use.

Strengths

  • Open weights model option for Whisper workflows.
  • Builders who want strong multilingual transcription quality with a lighter checkpoint for production-style speech pipelines.
  • The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.

Limitations

  • English only, which is the trade that buys the speed: it cannot transcribe or translate other languages at all, where every OpenAI Whisper checkpoint covers 99. Its licence is MIT rather than Apache 2.0.
  • ~1.5 GB in fp16, comparable to whisper-medium, but decoding is far cheaper because only two decoder layers run per token.
  • Context window and limits: 30-second audio windows · 448-token cap per window.
  • Verify the exact model card, provider docs, license, and serving support before production use.

Local workflow notes

The strongest choice for high-volume English transcription — podcast archives, meeting corpora, caption backfills — where throughput per GPU-hour is the constraint.

Local runtimes: Transformers, faster-whisper, whisper.cpp

Platforms: Windows, macOS, Linux

Transcription spec

Memory~1.5 GB in fp16 (756M parameters) · ~0.76 GB in int8Decoder depth2 layersLanguagesEnglish onlyAudio window30-second audio windows · 448-token cap per window

A distillation of large-v3 down to a 2-layer decoder — shallower than large-v3-turbo's 4 — with the 32-layer encoder kept intact. The fastest decoder on this page by a clear margin.

Sources to verify

Related resources

Continue with model source notes, local tools, and implementation guides related to this model.

Hardware~1.5 GB in fp16 (756M parameters)RuntimeTransformers, faster-whisper, whisper.cppContext30-second audio windows · 448-token cap per windowLast updated2026
Exact model card →

Model ecosystem connections

Use these next-step links to move from this profile into related tools, comparisons, guides, stacks, and curated shortlists.