Comparison of Open-Source Embedding Models
In this post, we’ll compare four of the top open-source embedding models that actually work in real-world pipelines. You’ll get:
- A breakdown of BGE, E5, Nomic, and MiniLM models, and when to use which
- Tradeoffs on accuracy, latency, and embedding speed
- A real benchmark on the BEIR TREC-COVID dataset, simulating RAG-style search
Why Go Open-Source?
Your embedding model is the backbone of a memory system or RAG pipeline. If you’re serious about optimization, transparency, or control, open-source models become the obvious choice.
First, they’re free to run and fine-tune. You can optimize them for your domain, deploy them wherever you want, and skip vendor lock-in. You’re also free to plug them into any system, like supermemory’s memory API, and scale up without being stuck in someone else’s pricing model or deployment timeline.
Second, open-source models let you see how things work under the hood. That means clearer debugging, better explainability, and smarter downstream usage when building vector pipelines.
Most importantly, they're catching up fast. Some open models now outperform proprietary ones in benchmarks, especially when you factor in retrieval accuracy and throughput; we’ll show you just how good these models are in the next section.
Best Open-Source Embedding Models
There are a lot of great options out there, but here are four open-source embedding models that stand out right now, especially for anyone building vector-based systems with retrieval, memory, or chat pipelines.
| Model | Size | Architecture | HuggingFace Link |
|---|---|---|---|
| BAAI/bge-base-en-v1.5 | 110M | BERT | https://huggingface.co/BAAI/bge-base-en-v1.5 |
| intfloat/e5-base-v2 | 110M | RoBERTa | https://huggingface.co/intfloat/e5-base-v2 |
| nomic-ai/nomic-embed-text-v1 | ~500M | GPT-style | https://huggingface.co/nomic-ai/nomic-embed-text-v1 |
| sentence-transformers/all-MiniLM-L6-v2 | 22M | MiniLM | https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 |
1. BAAI/bge-base-en-v1.5
A modern BERT-based model fine-tuned on dense retrieval tasks with contrastive learning and hard negatives. It supports both symmetric and asymmetric retrieval out of the box, and works well for reranking too.
Why choose it?
It's state-of-the-art on MTEB for English, super easy to plug into RAG systems, and supports query rewriting via prefixes like "Represent this sentence for retrieval". It’s widely used for academic and production search systems alike.
Disadvantages
While fast, it’s not the lightest model, and performance can drop when used on noisy or multilingual data. It also requires some pre-processing tweaks, like prefix prompting, to work optimally.
What’s under the hood?
- Architecture: Built on top of a BERT-style dual‑encoder, enabling super‑fast similarity search via FAISS-style vector lookup.
- Contrastive training with hard negatives: During fine-tuning, employs hard negative mining, sharpening its ability to rank relevant content highly.
- Instruction‑based prefix tuning: Fine-tuned to respond to prompts like "Represent this sentence for searching relevant passages:".
2. intfloat/e5-base-v2
Built on RoBERTa, this model performs well across tasks like search, reranking, and classification.
Why choose it?
It’s one of the most balanced models, with competitive accuracy, low latency, and robust across domains.
Disadvantages
For top performance, manage token length and truncation carefully. It may underperform slightly in some open-domain retrieval tasks.
What’s under the hood?
- Architecture: A RoBERTa base follows a bi-encoder architecture with shared Transformer encoder processing all text.
- Data Curation: Uses a large-scale, high-quality dataset (~270 million pairs) mined from various sources.
- Supervised Fine-Tuning with Labeled Data: Refined on smaller, labeled datasets to inject human-labeled nuance and relevance.
3. nomic-ai/nomic-embed-text-v1
This GPT-style embedding model was trained with a focus on high coverage and generalization.
Why choose it?
Excellent for large-scale search and memory systems.
Disadvantages
It is heavier and slower to embed compared to smaller models.
What’s under the hood?
- Custom long-context BERT backbone: Trained to support up to 8,192-token context.
- Multi-stage contrastive training (~235M text pairs): Refines on a high-quality dataset using contrastive learning.
- Instruction prefixes for task specialization: Supports multiple embedding roles.
4. sentence-transformers/all-MiniLM-L6-v2
This lightweight model is fast and resource-efficient.
Why choose it?
Blazing fast, low-resource, and easy to deploy.
Disadvantages
Not state-of-the-art in terms of retrieval accuracy, especially for complex tasks.
What’s under the hood?
- Lightweight MiniLM architecture: Based on a 6-layer MiniLM encoder.
- Optimized for short text (≈128–256 tokens): Trained with sequence length around 128 token pieces.
- Balances speed and quality: Fast performance, ideal for low-latency applications.
Benchmarking These Models
To evaluate the four models, we ran a benchmark using the BEIR TREC-COVID dataset, a popular benchmark for information retrieval.
Benchmarking Setup
Models Tested:
Benchmark Results
| Model | Embedding Time (ms/1K tokens) | Latency (Query → Retrieve) | Top-5 Retrieval Accuracy |
|---|---|---|---|
| MiniLM-L6-v2 | 14.7 | 68 ms | 78.1% |
| E5-Base-v2 | 20.2 | 79 ms | 83.5% |
| BGE-Base-v1.5 | 22.5 | 82 ms | 84.7% |
| Nomic Embed v1 | 41.9 | 110 ms | 86.2% |
Compute Cost Tradeoffs
| Model | GPU Memory Usage | Embedding Speed | Deployment Cost (Est.) |
|---|---|---|---|
| MiniLM-L6-v2 | ~1.2 GB | Very Fast | Low (Edge-compatible) |
| E5-Base-v2 | ~2.0 GB | Fast | Moderate |
| BGE-Base-v1.5 | ~2.1 GB | Medium | Moderate |
| Nomic Embed v1 | ~4.8 GB | Slow | High (GPU-dependent) |
Conclusion
Picking an embedding model directly shapes how helpful your LLMs can be. The benchmarks show that smaller models are fantastic when speed matters. Mid-size options offer a nice balance of power and efficiency. For accuracy-critical systems, larger models may be worth the investment.