Text Embeddings · Retrieval-Augmented Generation · Language Models

Sentence-BERT vs SimCSE vs E5: Same-Harness STS Numbers

Sentence-BERT cut 10k-sentence similarity from 65 hours to 5 seconds; SimCSE turned dropout into free training pairs for 76.25 unsupervised and 81.57 supervised STS; E5 beat BM25 zero-shot across 56 retrieval datasets.

Sentence-BERT vs SimCSE vs E5: Same-Harness STS Numbers

The one-line answer

These three papers are three answers to three different bottlenecks in sentence embeddings, published three years apart, and they barely compete on the same benchmark. Sentence-BERT (2019) solved an engineering problem: cross-encoder BERT scores every pair jointly, so finding the most similar pair among 10,000 sentences costs about 50 million inference computations, roughly 65 hours on a V100, while SBERT computes 10,000 embeddings in about 5 seconds and then compares them with cosine similarity. SimCSE (2021) solved a data problem: it uses dropout, the randomness already inside the Transformer, to create two slightly different views of the same sentence as a positive pair, which lets contrastive learning run with no labeled pairs at all and pushes unsupervised BERT-base from a 72.05 to a 76.25 averaged Spearman correlation on STS. E5 (2022) solved a scope problem: it pretrains one contrastive model on a large collection of weakly-supervised text pairs and evaluates it on 56 datasets spanning retrieval, clustering, and classification, becoming the first model family to beat the BM25 baseline on the BEIR retrieval benchmark in zero-shot mode without any labeled data.

What each method actually changed

Sentence-BERT’s change was architectural and economic. Vanilla BERT reads two sentences together and outputs a similarity score, which is accurate but quadratic: comparing n sentences means n-squared forward passes. SBERT fine-tunes BERT as a siamese network so that each sentence is encoded independently into a fixed-size vector, and similarity becomes a dot product between vectors. The paper’s memorable number is not an accuracy figure but the cost table: 65 hours of V100 time for the cross-encoder against roughly 5 seconds of embedding plus 0.01 seconds of cosine computation for the same 10,000-sentence task. Everything downstream of SBERT, including the embedding store at the heart of modern retrieval stacks, inherits this asymmetry.

SimCSE’s change was the definition of a positive pair. Contrastive learning needs two views of the same content, and earlier work manufactured them with crop, word deletion, or synonym replacement. The SimCSE authors noticed that passing the identical sentence through the same encoder twice already produces two different embeddings, because dropout randomly masks a different subset of activations each time. Training with these dropout pairs as positives beats every handcrafted augmentation on the STS-B development set. For the supervised variant, the paper reuses the NLI datasets (entailment as positive, contradiction as negative), which had been established as a strong recipe by SBERT two years earlier.

E5’s change was scale plus evaluation scope. Where SBERT and SimCSE optimize for semantic textual similarity, a narrow sentence-pair task, E5 targets general-purpose embeddings that must handle retrieval, clustering, and classification in one model. It pretrains with a simple contrastive recipe, in-batch negatives and a large batch, on a curated collection of weakly-supervised text pairs assembled from web sources, then evaluates on 56 datasets drawn from the MTEB and BEIR benchmark suites. The headline claims are about generality: zero-shot E5 is the first embedding model to outperform BM25 on BEIR retrieval without labeled data, and fine-tuned E5 reaches the best results on MTEB among models of its size class, beating existing embedding models with 40 times more parameters.

Key numbers

MeasurementSentence-BERTSimCSEE5Setting and sourceSame harness?
Most similar pair among 10,000 sentences, runtime~5 s (embed) + 0.01 s (cosine)n/an/aSBERT paper Sec. 1, V100Cross-encoder BERT in same passage: ~65 h / ~50M inference computations
STS average Spearman, BERT-base family, unsupervisedn/a76.25n/aSimCSE paper Table 1, 7 STS tasksYes: BERT-base scores 72.05 in the same table
STS average Spearman, BERT-base family, supervised (NLI)79.39 (SBERT-base row)81.57n/aSimCSE paper Table 2Yes: SBERT evaluated by the SimCSE authors in the same table
STS average Spearman, models not trained for STS77.03 (SBERT-NLI-base) / 79.23 (SBERT-NLI-large); raw BERT embeddings 46.35n/an/aSBERT paper Table 1SBERT-internal comparison
STS-B Spearman, fine-tuned directly on STS-B84.67 (SBERT-STSb-base); BERT-STSb-base 84.3082.5 (dev set, dropout pairs)n/aSBERT paper Table 1; SimCSE paper Table 1 (dev)Different settings; read the caveat below
Supervised SimCSE, RoBERTa-largen/a83.76n/aSimCSE paper Table 2SimCSE-internal scaling
Zero-shot retrieval, BEIRn/an/aOutperforms BM25 (first model family to do so without labeled data)E5 paper abstract and results, 56 datasets from MTEB + BEIRE5-internal suite; BM25 as the lexical reference
Fine-tuned retrieval/clustering/classification, MTEBn/an/aBest-in-class among compared models, beating models with 40x more parametersE5 paperE5-internal suite

Read the table column by column, because the three papers report on different tasks. The only strictly same-harness cross-method rows are the two middle SimCSE tables: their Table 1 pits unsupervised SimCSE-BERT-base against a BERT-base baseline they evaluated themselves, and their Table 2 pits it against an SBERT-base they also evaluated themselves. Everything else is a claim from one paper placed next to another’s, which is the honest way to read any embedding-model comparison on the internet.

SBERT’s real contribution was cost, and its own table admits it

The most instructive row in the SBERT paper is the one nobody quotes. When both BERT and SBERT are fine-tuned directly on the STS-B training set, the gap nearly vanishes: SBERT-STSb-base scores 84.67 Spearman and BERT-STSb-base scores 84.30. A cross-encoder, given the same labeled data, matches the dual encoder on the benchmark that the dual encoder was designed for. What it cannot do is run in 5 seconds over 10,000 sentences, because each comparison still requires a joint forward pass.

This is why the 65-hours-to-5-seconds figure is the actual contribution. Semantic search, deduplication, clustering, and retrieval at scale are inference-shape problems before they are accuracy problems. SBERT made the embedding-vector interface, encode once, compare forever, the default architecture for these tasks, and the two later papers in this comparison are best understood as attempts to raise the quality of what that interface returns.

Why dropout works as an augmentation

SimCSE’s ablation table on the STS-B development set quantifies how fragile handcrafted augmentations are for sentence semantics. Keeping the dropout-pair objective as-is scores 82.5; deleting one word drops it to 75.9; cropping to 90% of the length drops it to 77.8; replacing words with an MLM-predicted token collapses it to 62.2. The pattern makes sense: deleting or cropping a word changes the sentence’s meaning, so the model learns to treat meaning-damaged pairs as identical, which corrupts the geometry. Dropout perturbs activations without touching the words, so the positive pair genuinely shares semantics.

The paper also connects this to a known pathology of BERT embeddings: anisotropy, the tendency of pretrained embeddings to occupy a narrow cone in vector space, which makes cosine similarities between unrelated sentences suspiciously high. The authors show theoretically and empirically that the contrastive objective with dropout positives flattens this distribution. That framing explains the unsupervised headline: BERT-base averaged 72.05 Spearman across seven STS tasks in the SimCSE evaluation, and unsupervised SimCSE raises it to 76.25, a gain the paper summarizes as 4.2 points over the previous best unsupervised result. Adding NLI supervision on top reaches 81.57, which is 2.2 points above the previous best supervised result and, more concretely for this comparison, 2.18 points above the SBERT-base row in the same table.

E5 moves the goalpost from similarity to retrieval

The evaluation regime is the real difference between E5 and its two predecessors. STS tasks score pairs of short, clean, similar-length sentences. Retrieval, clustering, and classification ask whether one vector serves a query against a million-document index, groups documents by topic, or feeds a linear head. E5 reports on 56 datasets from the MTEB and BEIR suites, and its two claims are about that breadth. First, in zero-shot mode, where the model is used exactly as pretrained, E5 outperforms BM25 on the BEIR retrieval benchmark, something no previous embedding model had done without task-labeled fine-tuning. Second, when fine-tuned, E5 obtains the best results on MTEB among the compared models, including models with 40 times more parameters.

Both claims should be read as of 2022, because the embedding-model field has moved quickly since. But the durable idea is methodological: weakly-supervised contrastive pretraining on a large, curated pair corpus is enough to make a modestly sized encoder competitive across dozens of tasks, without per-task architecture changes. That recipe, a bidirectional encoder plus in-batch-negative contrastive loss plus large batch, is the template most open embedding models still follow.

When to use which

  • You need embeddings from a BERT-size model you already have, for similarity or clustering, and you have no labeled pairs: unsupervised SimCSE. It is a few lines on top of a pretrained encoder, and 76.25 averaged Spearman is a strong starting point.
  • You have NLI-style labeled pairs or can reuse SNLI and MultiNLI: supervised SimCSE, which reached 81.57 averaged Spearman with BERT-base in its own evaluation, ahead of the SBERT-base baseline in the same table.
  • You are building retrieval, clustering, or classification over web-scale text and need one general-purpose embedding: E5’s regime. Its contribution is the breadth of the 56-dataset evaluation and the zero-shot BM25 result, not a win on STS.
  • You are choosing between a cross-encoder and an embedding model for a similarity product: remember the SBERT table. If accuracy on a small labeled set is all that matters and joint encoding is affordable, the cross-encoder ties. If you must encode once and compare millions of times, embeddings are the only shape that works, and the quality question is which of the three regimes above fits your data.

Limits and open questions

All three papers evaluate on English benchmarks, and the numbers above do not transfer automatically to other languages or to domains like legal, medical, or code text. The same-harness rows are limited to the SimCSE evaluation, which is also the only paper here that re-scored its competitor’s models; the SBERT and E5 numbers come from their own protocols and their own base encoders, so cross-paper deltas smaller than a couple of points are noise. The field has also moved on: instruction-tuned and API embedding models now dominate the leaderboards these papers helped define, and anisotropy mitigations more sophisticated than dropout pairs exist. Treat these three papers as the reference points for why embeddings are fast, why contrastive objectives fix their geometry, and why retrieval-grade evaluation needs more than STS.

FAQ

Which is better on STS, Sentence-BERT or SimCSE?

Under the evaluation the SimCSE authors ran, SimCSE: supervised SimCSE-BERT-base averaged 81.57 Spearman across seven STS tasks, against 79.39 for the SBERT-base row in the same table. Unsupervised SimCSE, which needs no labeled pairs at all, reached 76.25 against BERT-base’s 72.05 in the same evaluation. SimCSE is the better choice when you can retrain; SBERT remains the historically important baseline whose models the SimCSE paper itself used as a supervised reference.

Why are raw BERT embeddings bad for cosine similarity?

Because of anisotropy: pretrained BERT embeddings occupy a narrow cone in vector space, so cosine similarity between unrelated sentences stays high and discriminative power collapses. The SBERT paper measured this directly, with raw BERT embeddings averaging 46.35 Spearman on STS tasks, far below even simple GloVe averages. SimCSE showed both theoretically and empirically that a contrastive objective with dropout positive pairs flattens the embedding distribution, which is a large part of where its unsupervised gain comes from.

Is SimCSE really unsupervised? Where do the positive pairs come from?

Yes. The positive pair is the same sentence passed through the same encoder twice. Dropout randomly masks a different subset of activations on each pass, so the two embeddings differ slightly but encode identical words. On the STS-B development set this beats explicit augmentations by wide margins: 82.5 for dropout pairs versus 75.9 for delete-one-word, 77.8 for 10% cropping, and 62.2 for MLM-based replacement.

What did Sentence-BERT actually contribute if fine-tuned BERT matches it on STS-B?

Speed at index time. When both are fine-tuned on STS-B, the accuracies nearly tie (84.67 versus 84.30). But the cross-encoder needs a joint forward pass per pair, about 65 hours for all pairs among 10,000 sentences on a V100, while SBERT encodes each sentence once, about 5 seconds total, and compares vectors with cosine similarity. SBERT’s contribution was making the embedding interface practical, not winning the benchmark.

When should I use E5 instead of SimCSE?

When your task is retrieval, clustering, or classification over general web text rather than sentence-pair similarity, and especially when you want zero-shot behavior. E5 was evaluated on 56 datasets from MTEB and BEIR, was the first embedding model family to beat BM25 on BEIR in zero-shot mode without labeled data, and when fine-tuned beat models with 40 times more parameters on MTEB. For a narrow STS-style similarity task with a BERT budget, the SimCSE numbers are the more direct evidence.