Model family

Microsoft / intfloatUpdated 2026EmbeddingRAGMultilingualSearch

E5 Models

E5 models are widely used for multilingual embeddings, semantic search, retrieval, low-overhead indexing, and RAG pipelines.

Best for

Embedding

Use this family hub to compare E5 variants for embedding workflows, then open the detail page for deeper deployment notes.

RAG

Use this family hub to compare E5 variants for rag workflows, then open the detail page for deeper deployment notes.

Multilingual

Use this family hub to compare E5 variants for multilingual workflows, then open the detail page for deeper deployment notes.

Search

Use this family hub to compare E5 variants for search workflows, then open the detail page for deeper deployment notes.

Variants

E5 models grouped by workflow

Embedding and reranking

EmbeddingRAGSemantic searchMultilingual

Multilingual E5 Large

Microsoft / intfloat · E5

Best for: Teams building multilingual retrieval, semantic search, and RAG pipelines.

Local: Runs locally for many embedding and semantic search prototypes on CPU or modest GPU hardware.

Details →
EmbeddingOpen weightsembeddingrag

e5-mistral-7b-instruct

Microsoft / intfloat · E5

Best for: Teams comparing larger instruction-tuned embedding models for retrieval quality when smaller E5 checkpoints are not enough.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

multilingual-e5-large-v2

Microsoft / intfloat · E5

Best for: Teams building multilingual search or RAG where one English-only embedding model is not enough.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

e5-large-v2

Microsoft / intfloat · E5

Best for: Teams that want a strong English retrieval baseline for search and RAG before moving to multilingual or larger embedding experiments.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

e5-base-v2

Microsoft / intfloat · E5

Best for: Builders who want a balanced English embedding model for search and RAG on moderate local or hosted infrastructure.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

e5-small-v2

Microsoft / intfloat · E5

Best for: Local or cost-sensitive embedding pipelines that need a practical English retrieval model without the overhead of larger E5 releases.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

e5-large

Microsoft / intfloat · E5

Best for: Legacy large embedding baseline

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

e5-base

Microsoft / intfloat · E5

Best for: Legacy base embedding baseline

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

e5-small

Microsoft / intfloat · E5

Best for: Legacy lightweight embedding baseline

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →

Compare

All E5 models in the directory

ModelTypeBest forLocal runner notesLicenseDetail
Multilingual E5 LargeEmbeddingTeams building multilingual retrieval, semantic search, and RAG pipelines.Runs locally for many embedding and semantic search prototypes on CPU or modest GPU hardware.MITOpen
e5-mistral-7b-instructEmbeddingTeams comparing larger instruction-tuned embedding models for retrieval quality when smaller E5 checkpoints are not enough.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
multilingual-e5-large-v2EmbeddingTeams building multilingual search or RAG where one English-only embedding model is not enough.Use the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
e5-large-v2EmbeddingTeams that want a strong English retrieval baseline for search and RAG before moving to multilingual or larger embedding experiments.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
e5-base-v2EmbeddingBuilders who want a balanced English embedding model for search and RAG on moderate local or hosted infrastructure.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
e5-small-v2EmbeddingLocal or cost-sensitive embedding pipelines that need a practical English retrieval model without the overhead of larger E5 releases.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
e5-largeEmbeddingLegacy large embedding baselineUse the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
e5-baseEmbeddingLegacy base embedding baselineUse the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
e5-smallEmbeddingLegacy lightweight embedding baselineUse the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen