Blog

/

Guides

/

Best Local Embedding Models for RAG: 7 Models Compared

Best Local Embedding Models for RAG: 7 Models Compared

Seven local embedding models compared on two English retrieval datasets. See how retrieval quality trades off against indexing speed, query latency, and memory.

Best Local Embedding Models for RAG: 7 Models Compared
Alex Shapiro
Alex Shapiro
Calendar icon

September 25, 2026

Table of Contents

Test scope: Seven local models, two English scientific corpora, and dense retrieval. These results do not establish multilingual or hybrid-search winners. See the methodology for the limits of comparison.

We tested seven local text embedding models on SciFact and NFCorpus, comparing retrieval quality, indexing speed, and memory use.

  • Qwen3-Embedding-8B recorded the highest nDCG@10 on both datasets, with the highest memory use and slowest indexing. EmbeddingGemma had the higher Recall@10.
  • EmbeddingGemma-300M came close to Qwen3-8B on nDCG@10 while using 0.96 GiB of peak VRAM, but had the slowest single-query latency in this comparison.
  • all-MiniLM-L6-v2 indexed fastest in this test at 1,074.7 chunks/s and produced the smallest native index.
  • Multilingual E5 Large supports 100 languages. These English-only tests do not establish a multilingual winner.

Models at a glance

Model specifications
ModelParametersDimensionsModel input limitWeight filesLicense
BAAI/bge-m3567.8M1,0248,1922.12 GiBMIT
nomic-ai/nomic-embed-text-v1.5136.7M7688,1920.51 GiBApache-2.0
Qwen/Qwen3-Embedding-0.6B595.8M1,02432k1.11 GiBApache-2.0
intfloat/multilingual-e5-large559.9M1,0245122.09 GiBMIT
google/embeddinggemma-300m300M class7682,0481.15 GiBGemma Terms
Qwen/Qwen3-Embedding-8B7.57B4,09632k14.10 GiBApache-2.0
sentence-transformers/all-MiniLM-L6-v222.7M3842560.08 GiBApache-2.0

Weight sizes refer to repository files, not runtime memory. EmbeddingGemma includes both projection layers. Qwen supports a 32k context; the tests capped it at 8,192 tokens. MiniLM's limit is measured in wordpieces.

Best Embedding Models: Quality, Speed, and Memory

The tables below separate retrieval quality from speed and memory. Qwen3-8B had the highest nDCG@10; EmbeddingGemma had the highest Recall@10 on both datasets.

Retrieval quality: higher is better
ModelSciFact nDCG@10SciFact Recall@10NFCorpus nDCG@10NFCorpus Recall@10
Qwen3-Embedding-8B0.79530.91830.40750.1968
EmbeddingGemma-300M0.78610.92020.39010.2013
Nomic Embed Text v1.50.71560.84110.34880.1736
Qwen3-Embedding-0.6B0.70110.83320.35670.1692
Multilingual E5 Large0.68520.79340.32990.1553
all-MiniLM-L6-v20.65290.82290.31830.1602
BGE-M30.65100.79510.31550.1518
BM25 baseline0.64020.77070.29690.1471
Speed, memory, and vector storage
ModelQuery p50 / p95Indexing chunks/sPeak VRAMSciFact index
Qwen3-Embedding-8B34.9 / 42.1 ms26.116.60 GiB152.5 MiB
EmbeddingGemma-300M43.0 / 46.6 ms315.00.96 GiB28.6 MiB
Nomic Embed Text v1.510.2 / 11.3 ms727.30.47 GiB28.6 MiB
Qwen3-Embedding-0.6B32.0 / 37.6 ms169.92.49 GiB38.1 MiB
Multilingual E5 Large18.5 / 20.1 ms434.31.29 GiB38.1 MiB
all-MiniLM-L6-v26.2 / 7.1 ms1,074.70.24 GiB14.3 MiB
BGE-M315.2 / 17.4 ms433.01.30 GiB38.1 MiB

nDCG@10 measures ranking quality. Recall@10 measures the share of relevant documents returned in the first ten results. Neither is an answer-accuracy percentage.

Using dense retrieval with native vector dimensions, each text chunk was converted into a single normalized vector and compared using cosine similarity without any post-retrieval reranking. The index size reflects the uncompressed float32 vector matrix required for the SciFact dataset, excluding any additional database overhead.

Query p50 and p95 are median and 95th-percentile latencies. The timing scope is not specified in the test report, so these figures should not be treated as end-to-end RAG response times.

Indexing throughput measures encoded chunks per second. Weight-file size, peak VRAM, and vector-index size measure different resources.

EmbeddingGemma-300M: High Recall, Low Memory

EmbeddingGemma-300M is Google's compact text embedding model, built from Gemma 3 for retrieval on phones, laptops, and desktops.

Model details
Model specificationEmbeddingGemma-300M
Languages100+ spoken languages
Query prompttask: search result | query:
Document prompttitle: none | text:
PoolingMean
Native vector size768 dimensions
Matryoshka support768, 512, 256, and 128 dimensions
Input limit2,048 tokens
Supported dtypefloat32 or bfloat16, not float16
LicenseGemma Terms

Official model card

EmbeddingGemma had the second-highest nDCG@10 and the highest Recall@10 in these tests, using 0.96 GiB of peak VRAM.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.7861 / 0.9202
NFCorpus nDCG@10 / Recall@100.3901 / 0.2013
Query latency, p50 / p9543.0 / 46.6 ms
Indexing speed315.0 chunks/s
Peak VRAM0.96 GiB
SciFact vector index28.6 MiB

However, it also had the highest query latency of the seven models.

Choose EmbeddingGemma when retrieval quality and low memory use matter more than single-query latency.

all-MiniLM-L6-v2: Fastest and Smallest

all-MiniLM-L6-v2 is a six-layer sentence-transformer trained on more than one billion text pairs.

It maps short sentences and paragraphs into compact dense vectors for semantic search, clustering, and similarity. Its small checkpoint and simple input format have made it a common baseline for local retrieval.

Model details
Model specificationall-MiniLM-L6-v2
LanguageEnglish
Query / document prefixNone / none
PoolingMean
Native vector size384 dimensions
Input limit256 wordpieces
LicenseApache-2.0

Official model card

MiniLM was the fastest model we measured. It also produced the smallest native index and used the least peak VRAM of the seven models.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.6529 / 0.8229
NFCorpus nDCG@10 / Recall@100.3183 / 0.1602
Query latency, p50 / p956.2 / 7.1 ms
Indexing speed1,074.7 chunks/s
Peak VRAM0.24 GiB
SciFact vector index14.3 MiB

However, when it comes to embedding quality, the model was substantially behind EmbeddingGemma and Qwen3-8B.

Documents longer than the model's 256-wordpiece input limit need chunking to avoid truncation.

Verdict: Choose MiniLM when indexing speed, low query latency, and a compact vector database matter the most.

Nomic Embed Text v1.5: Fast Midrange Option

Nomic Embed Text v1.5 is an English model designed for retrieval, classification, and clustering that supports Matryoshka embeddings.

Follow Nomic's official sequence: layer-normalize the full vector, truncate to the target dimension, then apply L2 normalization. This reduces vector storage; it does not reduce the model's parameter count.

Model details
Model specificationNomic Embed Text v1.5
LanguageEnglish
Query prefixsearch_query:
Document prefixsearch_document:
PoolingMean
Native vector size768 dimensions
Supported MRL sizes512, 256, 128, and 64 dimensions
Input limit8,192 tokens
LicenseApache-2.0

Official model card

In our test, Nomic was the second-fastest model for both query latency and bulk indexing, just behind MiniLM.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.7156 / 0.8411
NFCorpus nDCG@10 / Recall@100.3488 / 0.1736
Query latency, p50 / p9510.2 / 11.3 ms
Indexing speed727.3 chunks/s
Peak VRAM0.47 GiB
SciFact vector index28.6 MiB

In our separate Matryoshka run on SciFact, cutting the vector size from 768 to 512 dimensions resulted in a minimal drop in nDCG (0.7179 full-size to 0.7111). Accuracy fell more noticeably at 256 dimensions (0.6857).

Verdict: Choose Nomic when low memory use, short query latency, and quick reindexing matter more than reaching the highest retrieval score.

Qwen3-Embedding-0.6B: A Compact Qwen Option

Qwen3-Embedding-0.6B is the smallest Qwen embedding model.

Model details
Model specificationQwen3-Embedding-0.6B
Languages100+, including programming languages
Query formatRegistered query prompt with task instruction
Document prefixNone
PoolingLast token with left padding
Vector size32 to 1,024 dimensions
Official context window32k tokens
Limit used in our shared test8,192 tokens
LicenseApache-2.0

Official model card

Qwen3-0.6B landed near the middle of our retrieval rankings. While it outperformed Nomic on NFCorpus nDCG, Nomic still retained more relevant documents in its top ten results.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.7011 / 0.8332
NFCorpus nDCG@10 / Recall@100.3567 / 0.1692
Query latency, p50 / p9532.0 / 37.6 ms
Indexing speed169.9 chunks/s
Peak VRAM2.49 GiB
SciFact vector index38.1 MiB

Nomic had lower query latency and faster indexing, while EmbeddingGemma had higher retrieval scores with less peak memory. Use a task instruction for Qwen queries, as recommended by its model card; documents do not need that instruction.

Verdict: Choose Qwen3-Embedding-0.6B if you need Qwen's multilingual and instruction-aware embedding stack in a smaller checkpoint.

Multilingual E5 Large: Retrieval Across 100 Languages

Multilingual E5 Large is a 24-layer XLM-RoBERTa-based model trained on multilingual text pairs. It is designed for asymmetric retrieval, where a short query searches a collection of longer passages.

Model details
Model specificationMultilingual E5 Large
Languages100
Query prefixquery:
Document prefixpassage:
PoolingMean
Native vector size1,024 dimensions
Input limit512 tokens
LicenseMIT

Official model card

Our results on the two English datasets:

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.6852 / 0.7934
NFCorpus nDCG@10 / Recall@100.3299 / 0.1553
Query latency, p50 / p9518.5 / 20.1 ms
Indexing speed434.3 chunks/s
Peak VRAM1.29 GiB
SciFact vector index38.1 MiB

Note that this model has a 512-token input limit, which makes chunking necessary for long documents.

Verdict: Choose Multilingual E5 Large when you need a retrieval model for multiple languages.

Qwen3-Embedding-8B: Highest nDCG@10 in These Tests

Qwen3-Embedding-8B is the largest checkpoint in the Qwen3 embedding series, expanding the vector size fourfold compared to the 0.6B variant.

Model details
Model specificationQwen3-Embedding-8B
Languages100+, including programming languages
Query formatRegistered query prompt with task instruction
Document prefixNone
PoolingLast token with left padding
Vector size32 to 4,096 dimensions
Official context window32k tokens
Limit used in our shared test8,192 tokens
LicenseApache-2.0

Official model card

Qwen3-Embedding-8B scored 0.0174 higher than EmbeddingGemma on NFCorpus nDCG@10 and 0.0092 higher on SciFact. EmbeddingGemma had higher Recall@10 on both. No statistical significance test is reported.

Qwen3-8B indexed 26.1 chunks/s, compared with 169.9 for Qwen3-0.6B and 1,074.7 for MiniLM. Its native 4,096-dimension vectors require four times the raw storage of 1,024-dimension vectors.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.7953 / 0.9183
NFCorpus nDCG@10 / Recall@100.4075 / 0.1968
Query latency, p50 / p9534.9 / 42.1 ms
Indexing speed26.1 chunks/s
Peak VRAM16.60 GiB
SciFact vector index152.5 MiB

In the separate Matryoshka run, Qwen3-8B scored 0.7960 at 4,096 dimensions, 0.7905 at 1,024, and 0.7672 at 256 on SciFact nDCG@10. Cutting to 1,024 dimensions reduces the raw vector matrix by 75%; it does not reduce model weights or their memory use.

Choose Qwen3-Embedding-8B when its higher nDCG@10 on your corpus justifies the extra memory and slower indexing.

BGE-M3: Dense, Sparse, and Multi-Vector Retrieval

BGE-M3 is a retrieval model from the Beijing Academy of Artificial Intelligence, designed to support several search strategies from one checkpoint, making it useful if you want to combine semantic and keyword search without maintaining separate embedding models.

Model details
Model specificationBGE-M3
Retrieval modesDense, sparse, and multi-vector
Languages100+
Query / document prefixNone / none
Pooling for dense vectorsCLS token
Native vector size1,024 dimensions
Input limit8,192 tokens
LicenseMIT

Official model card

This comparison covers dense single-vector retrieval. Sparse, hybrid, and multi-vector modes were not evaluated.

Benchmark results
RTX A6000 testResult
SciFact nDCG@10 / Recall@100.6510 / 0.7951
NFCorpus nDCG@10 / Recall@100.3155 / 0.1518
Query latency, p50 / p9515.2 / 17.4 ms
Indexing speed433.0 chunks/s
Peak VRAM1.30 GiB
SciFact vector index38.1 MiB

Verdict: Choose BGE-M3 when you plan to use multilingual, sparse, or multi-vector retrieval.

How We Tested Retrieval Quality, Speed, and Memory

We used the full SciFact and NFCorpus corpora from BEIR, with the official test queries and relevance judgments.

  • SciFact contains 5,183 documents and 300 test queries.
  • NFCorpus contains 3,633 documents and 323 test queries.

The test used shared chunks capped at roughly 250 words. This is not a guarantee of equal tokenized input: 250 words can exceed MiniLM's 256-wordpiece limit. The report does not specify truncation rates or how chunk scores were combined into document scores, which limits interpretation of the ranking differences.

The models used their native prompts, pooling strategies, and tokenization, primarily in bfloat16 on one NVIDIA RTX A6000. Exact per-model dtype, batch size, library versions, and model revisions are not recorded in the report.

Retrieval quality was measured using standard ranking metrics (nDCG@10 and Recall@10), with median and 95th-percentile query latencies benchmarked over 100 warm runs.

Recall@10 depends on the number of relevant documents per query, so its values should not be compared directly between SciFact and NFCorpus. The BM25 baseline in this comparison scored 0.6402 on SciFact nDCG@10.

Small differences in nDCG require a paired analysis across queries before they can support claims of statistical significance. Repeat-run variation alone does not establish a universal significance threshold.

How to Choose an Embedding Model for Your RAG Pipeline

Where embeddings fit in a RAG pipeline.
The document collection and query must use the same embedding space. The generator uses retrieved passages to write an answer.

1. Start with your memory and latency limits

EmbeddingGemma is a candidate for strong retrieval with low memory use. MiniLM and Nomic had lower query latency in these measurements. A smaller checkpoint does not automatically produce faster queries.

2. Check retrieval errors before choosing a larger model

Inspect missed documents first. Check token truncation, query and document prompts, and support for your target language. Then compare another embedding model, keyword or hybrid retrieval, or a reranker if relevant documents are already present but ranked too low.

3. Optimize vector dimensions based on database constraints

If index storage or RAM becomes a bottleneck, use Matryoshka-capable models to truncate vector lengths.

Local vs API Embeddings

Local embedding models keep document text on hardware you control. You can pin an exact checkpoint, tune batching, and calculate the cost of a full reindex. This is useful for private corpora and predictable high-volume processing. Local does not mean free: you still pay for hardware, electricity, engineering time, and model monitoring.

An embedding API removes model serving and usually scales more easily. It can be the simpler choice for a small team or an uneven workload. The tradeoffs are per-token or per-request cost, data leaving your system, provider limits, and version management.

API response time includes network travel, queueing, and provider-side batching. Compare the complete pipelines you will operate, rather than treating an API request and a local GPU measurement as equivalent.

The generator can still be local whichever embedding route you choose. See how to run LLMs locally for the answer-generation side of a RAG system. For format and runtime choices, see our GGUF guide and local LLM apps.

How To Run These Models Locally

You can run these embedding models locally with Python and Sentence Transformers. Here's how:

Install Sentence Transformers:

pip install -U sentence-transformers

Save the following code as search.py:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B")

documents = [
    "Refunds are available within 30 days of purchase.",
    "The Pro plan includes ten team seats.",
    "Invoices can be downloaded from Billing settings.",
]
query = "Where can I get my invoice?"

document_vectors = model.encode_document(
    documents,
    normalize_embeddings=True,
)
query_vector = model.encode_query(
    query,
    normalize_embeddings=True,
)

scores = model.similarity(query_vector, document_vectors)[0]
top_results = scores.topk(k=2)

for score, index in zip(top_results.values, top_results.indices):
    print(f"{score.item():.3f}  {documents[index.item()]}")

Run it:

python search.py

The model downloads automatically on the first run. Sentence Transformers uses an available GPU when supported and otherwise runs on the CPU.

To test another model, use its supported loader, dtype, and query/document prompts. Re-encode every document and query when switching models: matching vector dimensions do not make two embedding spaces compatible.

FAQ

What is the best embedding model for RAG?

Qwen3-Embedding-8B had the highest nDCG@10 in the SciFact and NFCorpus results. EmbeddingGemma had higher Recall@10, lower peak VRAM, and faster indexing, but higher query latency. Choose using your own corpus and latency limits.

Do larger embedding models always perform better?

That's not always the case, for example, the 300M-class EmbeddingGemma model performs comparably to the 7.57B-parameter Qwen3 model. Key factors influencing retrieval quality include architecture design, training dataset diversity, retrieval prompt structure, pooling methods, and corpus characteristics.

Can I run embeddings locally?

Yes. These models provide downloadable weights for local inference under their respective licenses. Several used less than 1 GiB of peak VRAM in these tests, although those measurements came from a 48 GB RTX A6000. EmbeddingGemma uses Gemma Terms rather than an Apache or MIT license.

Do I need a reranker for RAG?

Not always. A reranker can improve the order of retrieved candidates, at the cost of extra latency and compute. It cannot recover a relevant document that the first retrieval stage did not return.

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Compare local AI video generators on quality, speed and VRAM, with RTX 4090 benchmarks, example clips and an Atomic Chat setup guide.

10/9/26

16 min

Best Local LLMs for 12GB VRAM in 2026

Best Local LLMs for 12GB VRAM in 2026

Eight local LLMs for 12GB VRAM: exact GGUF files, RTX 3080 Ti speed and memory figures, plus Snake and physics tests on an RTX 4070.

9/30/26

15 min

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 Alternatives Compared

Compare Claude Sonnet 5.5 with Opus, Fable and GPT-6 Astra, then choose a local Qwen, Ornith or Bonsai model for your hardware.

9/29/26

13 min

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Can you run Jev 1.13 locally? Learn how it works, see our Jev vs Laya Tetris demo, and set up Laya as an independent local alternative.

9/25/26

12 min