Comparison of Open-Source Embedding Models

In this post, we’ll compare four of the top open-source embedding models that actually work in real-world pipelines. You’ll get:

  • A breakdown of BGE, E5, Nomic, and MiniLM models, and when to use which
  • Tradeoffs on accuracy, latency, and embedding speed
  • A real benchmark on the BEIR TREC-COVID dataset, simulating RAG-style search

Why Go Open-Source?

Your embedding model is the backbone of a memory system or RAG pipeline. If you’re serious about optimization, transparency, or control, open-source models become the obvious choice.

First, they’re free to run and fine-tune. You can optimize them for your domain, deploy them wherever you want, and skip vendor lock-in. You’re also free to plug them into any system, like supermemory’s memory API, and scale up without being stuck in someone else’s pricing model or deployment timeline.

Second, open-source models let you see how things work under the hood. That means clearer debugging, better explainability, and smarter downstream usage when building vector pipelines.

Most importantly, they're catching up fast. Some open models now outperform proprietary ones in benchmarks, especially when you factor in retrieval accuracy and throughput; we’ll show you just how good these models are in the next section.

Best Open-Source Embedding Models

There are a lot of great options out there, but here are four open-source embedding models that stand out right now, especially for anyone building vector-based systems with retrieval, memory, or chat pipelines.

Model Size Architecture HuggingFace Link
BAAI/bge-base-en-v1.5 110M BERT https://huggingface.co/BAAI/bge-base-en-v1.5
intfloat/e5-base-v2 110M RoBERTa https://huggingface.co/intfloat/e5-base-v2
nomic-ai/nomic-embed-text-v1 ~500M GPT-style https://huggingface.co/nomic-ai/nomic-embed-text-v1
sentence-transformers/all-MiniLM-L6-v2 22M MiniLM https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2

1. BAAI/bge-base-en-v1.5

A modern BERT-based model fine-tuned on dense retrieval tasks with contrastive learning and hard negatives. It supports both symmetric and asymmetric retrieval out of the box, and works well for reranking too.

Why choose it?

It's state-of-the-art on MTEB for English, super easy to plug into RAG systems, and supports query rewriting via prefixes like "Represent this sentence for retrieval". It’s widely used for academic and production search systems alike.

Disadvantages

While fast, it’s not the lightest model, and performance can drop when used on noisy or multilingual data. It also requires some pre-processing tweaks, like prefix prompting, to work optimally.

What’s under the hood?

  • Architecture: Built on top of a BERT-style dual‑encoder, enabling super‑fast similarity search via FAISS-style vector lookup.
  • Contrastive training with hard negatives: During fine-tuning, employs hard negative mining, sharpening its ability to rank relevant content highly.
  • Instruction‑based prefix tuning: Fine-tuned to respond to prompts like "Represent this sentence for searching relevant passages:".

2. intfloat/e5-base-v2

Built on RoBERTa, this model performs well across tasks like search, reranking, and classification.

Why choose it?

It’s one of the most balanced models, with competitive accuracy, low latency, and robust across domains.

Disadvantages

For top performance, manage token length and truncation carefully. It may underperform slightly in some open-domain retrieval tasks.

What’s under the hood?

  • Architecture: A RoBERTa base follows a bi-encoder architecture with shared Transformer encoder processing all text.
  • Data Curation: Uses a large-scale, high-quality dataset (~270 million pairs) mined from various sources.
  • Supervised Fine-Tuning with Labeled Data: Refined on smaller, labeled datasets to inject human-labeled nuance and relevance.

3. nomic-ai/nomic-embed-text-v1

This GPT-style embedding model was trained with a focus on high coverage and generalization.

Why choose it?

Excellent for large-scale search and memory systems.

Disadvantages

It is heavier and slower to embed compared to smaller models.

What’s under the hood?

  • Custom long-context BERT backbone: Trained to support up to 8,192-token context.
  • Multi-stage contrastive training (~235M text pairs): Refines on a high-quality dataset using contrastive learning.
  • Instruction prefixes for task specialization: Supports multiple embedding roles.

4. sentence-transformers/all-MiniLM-L6-v2

This lightweight model is fast and resource-efficient.

Why choose it?

Blazing fast, low-resource, and easy to deploy.

Disadvantages

Not state-of-the-art in terms of retrieval accuracy, especially for complex tasks.

What’s under the hood?

  • Lightweight MiniLM architecture: Based on a 6-layer MiniLM encoder.
  • Optimized for short text (≈128–256 tokens): Trained with sequence length around 128 token pieces.
  • Balances speed and quality: Fast performance, ideal for low-latency applications.

Benchmarking These Models

To evaluate the four models, we ran a benchmark using the BEIR TREC-COVID dataset, a popular benchmark for information retrieval.

Benchmarking Setup

Models Tested:

Benchmark Results

Model Embedding Time (ms/1K tokens) Latency (Query → Retrieve) Top-5 Retrieval Accuracy
MiniLM-L6-v2 14.7 68 ms 78.1%
E5-Base-v2 20.2 79 ms 83.5%
BGE-Base-v1.5 22.5 82 ms 84.7%
Nomic Embed v1 41.9 110 ms 86.2%

Compute Cost Tradeoffs

Model GPU Memory Usage Embedding Speed Deployment Cost (Est.)
MiniLM-L6-v2 ~1.2 GB Very Fast Low (Edge-compatible)
E5-Base-v2 ~2.0 GB Fast Moderate
BGE-Base-v1.5 ~2.1 GB Medium Moderate
Nomic Embed v1 ~4.8 GB Slow High (GPU-dependent)

Conclusion

Picking an embedding model directly shapes how helpful your LLMs can be. The benchmarks show that smaller models are fantastic when speed matters. Mid-size options offer a nice balance of power and efficiency. For accuracy-critical systems, larger models may be worth the investment.