Model family

BAAIUpdated 2026EmbeddingRerankingRAGSearch

BGE Models

BGE models cover embeddings, reranking, retrieval, semantic search, and vector database workflows for RAG builders.

Best for

Embedding

Use this family hub to compare BGE variants for embedding workflows, then open the detail page for deeper deployment notes.

Reranking

Use this family hub to compare BGE variants for reranking workflows, then open the detail page for deeper deployment notes.

RAG

Use this family hub to compare BGE variants for rag workflows, then open the detail page for deeper deployment notes.

Search

Use this family hub to compare BGE variants for search workflows, then open the detail page for deeper deployment notes.

Variants

BGE models grouped by workflow

Embedding and reranking

RerankingRAGRetrievalSearch

BGE Reranker v2 M3

BAAI · BGE

Best for: RAG builders who need a practical reranker after Qdrant, Chroma, pgvector, or other vector search.

Local: Commonly used in local RAG stacks as a reranking step after vector search.

Details →
EmbeddingOpen weightsembeddingrag

bge-m3

BAAI · BGE

Best for: RAG teams that want one embedding model for multilingual search, long documents, and retrieval experiments beyond plain dense vectors.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

bge-large-en-v1.5

BAAI · BGE

Best for: English retrieval stacks that want a stronger embedding baseline for search and RAG before moving to multilingual models.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

bge-base-en-v1.5

BAAI · BGE

Best for: Teams building English search or RAG systems on modest GPUs without dropping to the smallest embedding tier.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weightsembeddingrag

bge-small-en-v1.5

BAAI · BGE

Best for: Lightweight English embedding pipelines where CPU or modest GPU serving matters more than chasing the largest model.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

bge-large-zh-v1.5

BAAI · BGE

Best for: Chinese embedding and retrieval workflows

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
RerankingOpen weightsrerankingrag

bge-reranker-large

BAAI · BGE

Best for: Teams that already retrieve a candidate set and want a stronger final ranking step before passing context to an LLM.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
RerankingOpen weightsrerankingrag

bge-reranker-base

BAAI · BGE

Best for: Builders who want a practical reranking step in RAG or search without adding a heavyweight cross-encoder to every request.

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →
EmbeddingOpen weights where releasedembeddingrag

bge-embedding-gemma2

BAAI · BGE

Best for: Gemma-based embedding experiments

Local: Use the exact checkpoint and quantization that matches your hardware and latency target.

Details →

Compare

All BGE models in the directory

ModelTypeBest forLocal runner notesLicenseDetail
BGE Reranker v2 M3RerankingRAG builders who need a practical reranker after Qdrant, Chroma, pgvector, or other vector search.Commonly used in local RAG stacks as a reranking step after vector search.Check model cardOpen
bge-m3EmbeddingRAG teams that want one embedding model for multilingual search, long documents, and retrieval experiments beyond plain dense vectors.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-large-en-v1.5EmbeddingEnglish retrieval stacks that want a stronger embedding baseline for search and RAG before moving to multilingual models.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-base-en-v1.5EmbeddingTeams building English search or RAG systems on modest GPUs without dropping to the smallest embedding tier.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-small-en-v1.5EmbeddingLightweight English embedding pipelines where CPU or modest GPU serving matters more than chasing the largest model.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-large-zh-v1.5EmbeddingChinese embedding and retrieval workflowsUse the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen
bge-reranker-largeRerankingTeams that already retrieve a candidate set and want a stronger final ranking step before passing context to an LLM.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-reranker-baseRerankingBuilders who want a practical reranking step in RAG or search without adding a heavyweight cross-encoder to every request.Use the exact checkpoint and quantization that matches your hardware and latency target.MITOpen
bge-embedding-gemma2EmbeddingGemma-based embedding experimentsUse the exact checkpoint and quantization that matches your hardware and latency target.Check exact model cardOpen