We tested seven local text embedding models on SciFact and NFCorpus, comparing retrieval quality, indexing speed, and memory use.
- Qwen3-Embedding-8B recorded the highest nDCG@10 on both datasets, with the highest memory use and slowest indexing. EmbeddingGemma had the higher Recall@10.
- EmbeddingGemma-300M came close to Qwen3-8B on nDCG@10 while using 0.96 GiB of peak VRAM, but had the slowest single-query latency in this comparison.
- all-MiniLM-L6-v2 indexed fastest in this test at 1,074.7 chunks/s and produced the smallest native index.
- Multilingual E5 Large supports 100 languages. These English-only tests do not establish a multilingual winner.
Models at a glance
| Model | Parameters | Dimensions | Model input limit | Weight files | License |
|---|---|---|---|---|---|
| BAAI/bge-m3 | 567.8M | 1,024 | 8,192 | 2.12 GiB | MIT |
| nomic-ai/nomic-embed-text-v1.5 | 136.7M | 768 | 8,192 | 0.51 GiB | Apache-2.0 |
| Qwen/Qwen3-Embedding-0.6B | 595.8M | 1,024 | 32k | 1.11 GiB | Apache-2.0 |
| intfloat/multilingual-e5-large | 559.9M | 1,024 | 512 | 2.09 GiB | MIT |
| google/embeddinggemma-300m | 300M class | 768 | 2,048 | 1.15 GiB | Gemma Terms |
| Qwen/Qwen3-Embedding-8B | 7.57B | 4,096 | 32k | 14.10 GiB | Apache-2.0 |
| sentence-transformers/all-MiniLM-L6-v2 | 22.7M | 384 | 256 | 0.08 GiB | Apache-2.0 |
Best Embedding Models: Quality, Speed, and Memory
The tables below separate retrieval quality from speed and memory. Qwen3-8B had the highest nDCG@10; EmbeddingGemma had the highest Recall@10 on both datasets.
| Model | SciFact nDCG@10 | SciFact Recall@10 | NFCorpus nDCG@10 | NFCorpus Recall@10 |
|---|---|---|---|---|
| Qwen3-Embedding-8B | 0.7953 | 0.9183 | 0.4075 | 0.1968 |
| EmbeddingGemma-300M | 0.7861 | 0.9202 | 0.3901 | 0.2013 |
| Nomic Embed Text v1.5 | 0.7156 | 0.8411 | 0.3488 | 0.1736 |
| Qwen3-Embedding-0.6B | 0.7011 | 0.8332 | 0.3567 | 0.1692 |
| Multilingual E5 Large | 0.6852 | 0.7934 | 0.3299 | 0.1553 |
| all-MiniLM-L6-v2 | 0.6529 | 0.8229 | 0.3183 | 0.1602 |
| BGE-M3 | 0.6510 | 0.7951 | 0.3155 | 0.1518 |
| BM25 baseline | 0.6402 | 0.7707 | 0.2969 | 0.1471 |
| Model | Query p50 / p95 | Indexing chunks/s | Peak VRAM | SciFact index |
|---|---|---|---|---|
| Qwen3-Embedding-8B | 34.9 / 42.1 ms | 26.1 | 16.60 GiB | 152.5 MiB |
| EmbeddingGemma-300M | 43.0 / 46.6 ms | 315.0 | 0.96 GiB | 28.6 MiB |
| Nomic Embed Text v1.5 | 10.2 / 11.3 ms | 727.3 | 0.47 GiB | 28.6 MiB |
| Qwen3-Embedding-0.6B | 32.0 / 37.6 ms | 169.9 | 2.49 GiB | 38.1 MiB |
| Multilingual E5 Large | 18.5 / 20.1 ms | 434.3 | 1.29 GiB | 38.1 MiB |
| all-MiniLM-L6-v2 | 6.2 / 7.1 ms | 1,074.7 | 0.24 GiB | 14.3 MiB |
| BGE-M3 | 15.2 / 17.4 ms | 433.0 | 1.30 GiB | 38.1 MiB |
Using dense retrieval with native vector dimensions, each text chunk was converted into a single normalized vector and compared using cosine similarity without any post-retrieval reranking. The index size reflects the uncompressed float32 vector matrix required for the SciFact dataset, excluding any additional database overhead.
Query p50 and p95 are median and 95th-percentile latencies. The timing scope is not specified in the test report, so these figures should not be treated as end-to-end RAG response times.
Indexing throughput measures encoded chunks per second. Weight-file size, peak VRAM, and vector-index size measure different resources.
EmbeddingGemma-300M: High Recall, Low Memory
EmbeddingGemma-300M is Google's compact text embedding model, built from Gemma 3 for retrieval on phones, laptops, and desktops.
| Model specification | EmbeddingGemma-300M |
|---|---|
| Languages | 100+ spoken languages |
| Query prompt | task: search result | query: |
| Document prompt | title: none | text: |
| Pooling | Mean |
| Native vector size | 768 dimensions |
| Matryoshka support | 768, 512, 256, and 128 dimensions |
| Input limit | 2,048 tokens |
| Supported dtype | float32 or bfloat16, not float16 |
| License | Gemma Terms |
EmbeddingGemma had the second-highest nDCG@10 and the highest Recall@10 in these tests, using 0.96 GiB of peak VRAM.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.7861 / 0.9202 |
| NFCorpus nDCG@10 / Recall@10 | 0.3901 / 0.2013 |
| Query latency, p50 / p95 | 43.0 / 46.6 ms |
| Indexing speed | 315.0 chunks/s |
| Peak VRAM | 0.96 GiB |
| SciFact vector index | 28.6 MiB |
However, it also had the highest query latency of the seven models.
Choose EmbeddingGemma when retrieval quality and low memory use matter more than single-query latency.
all-MiniLM-L6-v2: Fastest and Smallest
all-MiniLM-L6-v2 is a six-layer sentence-transformer trained on more than one billion text pairs.
It maps short sentences and paragraphs into compact dense vectors for semantic search, clustering, and similarity. Its small checkpoint and simple input format have made it a common baseline for local retrieval.
| Model specification | all-MiniLM-L6-v2 |
|---|---|
| Language | English |
| Query / document prefix | None / none |
| Pooling | Mean |
| Native vector size | 384 dimensions |
| Input limit | 256 wordpieces |
| License | Apache-2.0 |
MiniLM was the fastest model we measured. It also produced the smallest native index and used the least peak VRAM of the seven models.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.6529 / 0.8229 |
| NFCorpus nDCG@10 / Recall@10 | 0.3183 / 0.1602 |
| Query latency, p50 / p95 | 6.2 / 7.1 ms |
| Indexing speed | 1,074.7 chunks/s |
| Peak VRAM | 0.24 GiB |
| SciFact vector index | 14.3 MiB |
However, when it comes to embedding quality, the model was substantially behind EmbeddingGemma and Qwen3-8B.
Documents longer than the model's 256-wordpiece input limit need chunking to avoid truncation.
Verdict: Choose MiniLM when indexing speed, low query latency, and a compact vector database matter the most.
Nomic Embed Text v1.5: Fast Midrange Option
Nomic Embed Text v1.5 is an English model designed for retrieval, classification, and clustering that supports Matryoshka embeddings.
Follow Nomic's official sequence: layer-normalize the full vector, truncate to the target dimension, then apply L2 normalization. This reduces vector storage; it does not reduce the model's parameter count.
| Model specification | Nomic Embed Text v1.5 |
|---|---|
| Language | English |
| Query prefix | search_query: |
| Document prefix | search_document: |
| Pooling | Mean |
| Native vector size | 768 dimensions |
| Supported MRL sizes | 512, 256, 128, and 64 dimensions |
| Input limit | 8,192 tokens |
| License | Apache-2.0 |
In our test, Nomic was the second-fastest model for both query latency and bulk indexing, just behind MiniLM.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.7156 / 0.8411 |
| NFCorpus nDCG@10 / Recall@10 | 0.3488 / 0.1736 |
| Query latency, p50 / p95 | 10.2 / 11.3 ms |
| Indexing speed | 727.3 chunks/s |
| Peak VRAM | 0.47 GiB |
| SciFact vector index | 28.6 MiB |
In our separate Matryoshka run on SciFact, cutting the vector size from 768 to 512 dimensions resulted in a minimal drop in nDCG (0.7179 full-size to 0.7111). Accuracy fell more noticeably at 256 dimensions (0.6857).
Verdict: Choose Nomic when low memory use, short query latency, and quick reindexing matter more than reaching the highest retrieval score.
Qwen3-Embedding-0.6B: A Compact Qwen Option
Qwen3-Embedding-0.6B is the smallest Qwen embedding model.
| Model specification | Qwen3-Embedding-0.6B |
|---|---|
| Languages | 100+, including programming languages |
| Query format | Registered query prompt with task instruction |
| Document prefix | None |
| Pooling | Last token with left padding |
| Vector size | 32 to 1,024 dimensions |
| Official context window | 32k tokens |
| Limit used in our shared test | 8,192 tokens |
| License | Apache-2.0 |
Qwen3-0.6B landed near the middle of our retrieval rankings. While it outperformed Nomic on NFCorpus nDCG, Nomic still retained more relevant documents in its top ten results.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.7011 / 0.8332 |
| NFCorpus nDCG@10 / Recall@10 | 0.3567 / 0.1692 |
| Query latency, p50 / p95 | 32.0 / 37.6 ms |
| Indexing speed | 169.9 chunks/s |
| Peak VRAM | 2.49 GiB |
| SciFact vector index | 38.1 MiB |
Nomic had lower query latency and faster indexing, while EmbeddingGemma had higher retrieval scores with less peak memory. Use a task instruction for Qwen queries, as recommended by its model card; documents do not need that instruction.
Verdict: Choose Qwen3-Embedding-0.6B if you need Qwen's multilingual and instruction-aware embedding stack in a smaller checkpoint.
Multilingual E5 Large: Retrieval Across 100 Languages
Multilingual E5 Large is a 24-layer XLM-RoBERTa-based model trained on multilingual text pairs. It is designed for asymmetric retrieval, where a short query searches a collection of longer passages.
| Model specification | Multilingual E5 Large |
|---|---|
| Languages | 100 |
| Query prefix | query: |
| Document prefix | passage: |
| Pooling | Mean |
| Native vector size | 1,024 dimensions |
| Input limit | 512 tokens |
| License | MIT |
Our results on the two English datasets:
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.6852 / 0.7934 |
| NFCorpus nDCG@10 / Recall@10 | 0.3299 / 0.1553 |
| Query latency, p50 / p95 | 18.5 / 20.1 ms |
| Indexing speed | 434.3 chunks/s |
| Peak VRAM | 1.29 GiB |
| SciFact vector index | 38.1 MiB |
Note that this model has a 512-token input limit, which makes chunking necessary for long documents.
Verdict: Choose Multilingual E5 Large when you need a retrieval model for multiple languages.
Qwen3-Embedding-8B: Highest nDCG@10 in These Tests
Qwen3-Embedding-8B is the largest checkpoint in the Qwen3 embedding series, expanding the vector size fourfold compared to the 0.6B variant.
| Model specification | Qwen3-Embedding-8B |
|---|---|
| Languages | 100+, including programming languages |
| Query format | Registered query prompt with task instruction |
| Document prefix | None |
| Pooling | Last token with left padding |
| Vector size | 32 to 4,096 dimensions |
| Official context window | 32k tokens |
| Limit used in our shared test | 8,192 tokens |
| License | Apache-2.0 |
Qwen3-Embedding-8B scored 0.0174 higher than EmbeddingGemma on NFCorpus nDCG@10 and 0.0092 higher on SciFact. EmbeddingGemma had higher Recall@10 on both. No statistical significance test is reported.
Qwen3-8B indexed 26.1 chunks/s, compared with 169.9 for Qwen3-0.6B and 1,074.7 for MiniLM. Its native 4,096-dimension vectors require four times the raw storage of 1,024-dimension vectors.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.7953 / 0.9183 |
| NFCorpus nDCG@10 / Recall@10 | 0.4075 / 0.1968 |
| Query latency, p50 / p95 | 34.9 / 42.1 ms |
| Indexing speed | 26.1 chunks/s |
| Peak VRAM | 16.60 GiB |
| SciFact vector index | 152.5 MiB |
In the separate Matryoshka run, Qwen3-8B scored 0.7960 at 4,096 dimensions, 0.7905 at 1,024, and 0.7672 at 256 on SciFact nDCG@10. Cutting to 1,024 dimensions reduces the raw vector matrix by 75%; it does not reduce model weights or their memory use.
Choose Qwen3-Embedding-8B when its higher nDCG@10 on your corpus justifies the extra memory and slower indexing.
BGE-M3: Dense, Sparse, and Multi-Vector Retrieval
BGE-M3 is a retrieval model from the Beijing Academy of Artificial Intelligence, designed to support several search strategies from one checkpoint, making it useful if you want to combine semantic and keyword search without maintaining separate embedding models.
| Model specification | BGE-M3 |
|---|---|
| Retrieval modes | Dense, sparse, and multi-vector |
| Languages | 100+ |
| Query / document prefix | None / none |
| Pooling for dense vectors | CLS token |
| Native vector size | 1,024 dimensions |
| Input limit | 8,192 tokens |
| License | MIT |
This comparison covers dense single-vector retrieval. Sparse, hybrid, and multi-vector modes were not evaluated.
| RTX A6000 test | Result |
|---|---|
| SciFact nDCG@10 / Recall@10 | 0.6510 / 0.7951 |
| NFCorpus nDCG@10 / Recall@10 | 0.3155 / 0.1518 |
| Query latency, p50 / p95 | 15.2 / 17.4 ms |
| Indexing speed | 433.0 chunks/s |
| Peak VRAM | 1.30 GiB |
| SciFact vector index | 38.1 MiB |
Verdict: Choose BGE-M3 when you plan to use multilingual, sparse, or multi-vector retrieval.
How We Tested Retrieval Quality, Speed, and Memory
We used the full SciFact and NFCorpus corpora from BEIR, with the official test queries and relevance judgments.
- SciFact contains 5,183 documents and 300 test queries.
- NFCorpus contains 3,633 documents and 323 test queries.
The test used shared chunks capped at roughly 250 words. This is not a guarantee of equal tokenized input: 250 words can exceed MiniLM's 256-wordpiece limit. The report does not specify truncation rates or how chunk scores were combined into document scores, which limits interpretation of the ranking differences.
The models used their native prompts, pooling strategies, and tokenization, primarily in bfloat16 on one NVIDIA RTX A6000. Exact per-model dtype, batch size, library versions, and model revisions are not recorded in the report.
Retrieval quality was measured using standard ranking metrics (nDCG@10 and Recall@10), with median and 95th-percentile query latencies benchmarked over 100 warm runs.
Recall@10 depends on the number of relevant documents per query, so its values should not be compared directly between SciFact and NFCorpus. The BM25 baseline in this comparison scored 0.6402 on SciFact nDCG@10.
Small differences in nDCG require a paired analysis across queries before they can support claims of statistical significance. Repeat-run variation alone does not establish a universal significance threshold.
How to Choose an Embedding Model for Your RAG Pipeline
1. Start with your memory and latency limits
EmbeddingGemma is a candidate for strong retrieval with low memory use. MiniLM and Nomic had lower query latency in these measurements. A smaller checkpoint does not automatically produce faster queries.
2. Check retrieval errors before choosing a larger model
Inspect missed documents first. Check token truncation, query and document prompts, and support for your target language. Then compare another embedding model, keyword or hybrid retrieval, or a reranker if relevant documents are already present but ranked too low.
3. Optimize vector dimensions based on database constraints
If index storage or RAM becomes a bottleneck, use Matryoshka-capable models to truncate vector lengths.
Local vs API Embeddings
Local embedding models keep document text on hardware you control. You can pin an exact checkpoint, tune batching, and calculate the cost of a full reindex. This is useful for private corpora and predictable high-volume processing. Local does not mean free: you still pay for hardware, electricity, engineering time, and model monitoring.
An embedding API removes model serving and usually scales more easily. It can be the simpler choice for a small team or an uneven workload. The tradeoffs are per-token or per-request cost, data leaving your system, provider limits, and version management.
API response time includes network travel, queueing, and provider-side batching. Compare the complete pipelines you will operate, rather than treating an API request and a local GPU measurement as equivalent.
The generator can still be local whichever embedding route you choose. See how to run LLMs locally for the answer-generation side of a RAG system. For format and runtime choices, see our GGUF guide and local LLM apps.
How To Run These Models Locally
You can run these embedding models locally with Python and Sentence Transformers. Here's how:
Install Sentence Transformers:
pip install -U sentence-transformers
Save the following code as search.py:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B")
documents = [
"Refunds are available within 30 days of purchase.",
"The Pro plan includes ten team seats.",
"Invoices can be downloaded from Billing settings.",
]
query = "Where can I get my invoice?"
document_vectors = model.encode_document(
documents,
normalize_embeddings=True,
)
query_vector = model.encode_query(
query,
normalize_embeddings=True,
)
scores = model.similarity(query_vector, document_vectors)[0]
top_results = scores.topk(k=2)
for score, index in zip(top_results.values, top_results.indices):
print(f"{score.item():.3f} {documents[index.item()]}")
Run it:
python search.py
The model downloads automatically on the first run. Sentence Transformers uses an available GPU when supported and otherwise runs on the CPU.
To test another model, use its supported loader, dtype, and query/document prompts. Re-encode every document and query when switching models: matching vector dimensions do not make two embedding spaces compatible.
FAQ
What is the best embedding model for RAG?
Qwen3-Embedding-8B had the highest nDCG@10 in the SciFact and NFCorpus results. EmbeddingGemma had higher Recall@10, lower peak VRAM, and faster indexing, but higher query latency. Choose using your own corpus and latency limits.
Do larger embedding models always perform better?
That's not always the case, for example, the 300M-class EmbeddingGemma model performs comparably to the 7.57B-parameter Qwen3 model. Key factors influencing retrieval quality include architecture design, training dataset diversity, retrieval prompt structure, pooling methods, and corpus characteristics.
Can I run embeddings locally?
Yes. These models provide downloadable weights for local inference under their respective licenses. Several used less than 1 GiB of peak VRAM in these tests, although those measurements came from a 48 GB RTX A6000. EmbeddingGemma uses Gemma Terms rather than an Apache or MIT license.
Do I need a reranker for RAG?
Not always. A reranker can improve the order of retrieved candidates, at the cost of extra latency and compute. It cannot recover a relevant document that the first retrieval stage did not return.

