Fine-Tuning, Distillation and Custom Models · Evaluation, hosted platforms, embeddings and cost · lesson 13 of 16 · 15 min
Fine-tuning embedding models for better retrieval
Why retrieval is often the real bottleneck
In RAG systems, the generator can only use what retrieval finds. Generic embedding models are trained on broad web data; they may not understand your product names, internal jargon, abbreviations or code-mixed queries like Roman Urdu. Fine-tuning an embedding model on your own query-passage pairs is often one of the highest-return, lowest-cost customizations: embedding models are small (tens to hundreds of millions of parameters, sometimes more), train in minutes to hours on one GPU, and improvements show up directly in retrieval metrics.
Before fine-tuning, try the cheaper steps: hybrid search (keyword + vector), better chunking, metadata filters and a reranker (a cross-encoder that rescores top results). Fine-tune embeddings when retrieval still misses relevant passages.
How contrastive fine-tuning works
You provide (anchor, positive) pairs: a query and a passage that answers it. The common loss, Multiple Negatives Ranking Loss (MNRL), treats every other positive in the batch as a negative: it pulls each query toward its passage and pushes it away from the others. Larger batches give more negatives, and hard negatives (passages that look relevant but are not) teach fine distinctions.
Data sources for pairs:
- Search logs: query → the document the user clicked and did not bounce from.
- Support tickets: customer question → the help article the agent linked.
- Synthetic: a model writes realistic questions for each passage (then filter).
Hands-on: Sentence Transformers training
# train_embed.py (pip install -U sentence-transformers datasets)
from datasets import load_dataset
from sentence_transformers import (SentenceTransformer, SentenceTransformerTrainer,
SentenceTransformerTrainingArguments, mine_hard_negatives)
from sentence_transformers.sentence_transformer.losses import MultipleNegativesRankingLoss
from sentence_transformers.sentence_transformer.evaluation import InformationRetrievalEvaluator
BASE = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2" # pick a multilingual model that fits your license and languages
model = SentenceTransformer(BASE)
pairs = load_dataset("json", data_files="pairs.jsonl", split="train") # {"anchor": "...", "positive": "..."}
split = pairs.train_test_split(test_size=0.1, seed=42)
# Optional: add hard negatives mined with the current model
train = mine_hard_negatives(split["train"], model, anchor_column_name="anchor", positive_column_name="positive",
num_negatives=1, range_min=10, range_max=50, sampling_strategy="top")
# Retrieval evaluator on the held-out split
queries = {str(i): q for i, q in enumerate(split["test"]["anchor"])}
corpus = {str(i): p for i, p in enumerate(split["test"]["positive"])}
relevant = {str(i): {str(i)} for i in range(len(split["test"]))}
evaluator = InformationRetrievalEvaluator(queries, corpus, relevant, name="heldout")
print("before:", evaluator(model))
args = SentenceTransformerTrainingArguments(
output_dir="out/embed-ft", num_train_epochs=1, per_device_train_batch_size=64,
learning_rate=2e-5, warmup_steps=50, batch_sampler="no_duplicates", # avoid duplicate texts acting as false negatives
eval_strategy="steps", eval_steps=100, save_strategy="steps", save_steps=100, report_to="none")
trainer = SentenceTransformerTrainer(model=model, args=args, train_dataset=train,
loss=MultipleNegativesRankingLoss(model), evaluator=evaluator)
trainer.train()
print("after:", evaluator(model))
model.save_pretrained("out/embed-ft/final")
Sentence Transformers reorganized some import paths in recent major versions; if an import fails, check the current documentation. The evaluator reports metrics such as Recall@k, MRR@k and nDCG@k: compare before and after.
A realistic evaluation uses a full corpus (all your passages, not just the test positives) and real queries; the snippet above is a quick start.
Operational consequences
- Re-index everything. Vectors from the old and new models are not comparable. Re-embed the whole corpus, ideally into a new index, then switch.
- Version your index with the embedding model name and revision.
- Query and document prompts: some embedding models expect task prefixes; keep them consistent between training and serving.
- Serving: embedding models are small enough to run on CPU for moderate traffic, or via inference servers that support embeddings (for example vLLM, TEI or Ollama for supported models).
Worked example: Roman Urdu help-center search
A Pakistani fintech's help center is written in English, but a large share of searches are in Roman Urdu ("bill payment nahi ho rahi"). A generic multilingual model found the right article in the top 5 for about half of such queries (illustrative). The team built 8,000 pairs from agent-linked articles in support chats, added 4,000 filtered synthetic Roman Urdu questions, mined hard negatives, and fine-tuned for one epoch on a single GPU in under an hour. Recall@5 on real held-out Roman Urdu queries rose substantially, and English queries did not regress. They re-indexed 2,300 articles into a new versioned index and switched traffic after a shadow comparison.
Pitfalls
- Evaluating only on the pairs' own passages instead of the full corpus.
- Forgetting to re-index (mixing old and new vectors silently breaks search).
- Duplicate or near-duplicate passages treated as negatives (false negatives).
- Choosing a base embedding model without checking its license and languages.
How to measure success
Recall@k and MRR on real held-out queries against the full corpus improve over the base model (and over a reranker-only baseline), with no regression on other languages, and a versioned index rollout.
Video lecture: Fine-tuning embedding models for better retrieval
Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.
- Embedding fine-tuning
- Analogy: training a librarian
- Why and when
- Contrastive training
- Pair sources
- Hands-on: train_embed.py
- Operations
- Evaluate properly
- Worked example: Roman Urdu search (illustrative)
- Simple example: electronics search
- Reranker or fine-tune?
- FAQ: tune embeddings, generator or both?
- Try this now
- Watch me do it (illustrative)
- Recap
Lecture transcript
Embedding fine-tuning
Here is a secret about RAG systems. When answers are wrong, the language model often gets the blame, but the real culprit is retrieval. The right passage never reached the model. And one of the cheapest, highest-return fixes is fine-tuning the embedding model on your own queries. In this lesson you will learn when to do it, how contrastive training works, a complete Sentence Transformers script, and the operational steps people forget.
Analogy: training a librarian
Here is an analogy. An embedding model is a librarian who shelves books by meaning. A general librarian is great with popular topics but gets confused by your company's jargon, maybe shelving refund reversal next to reverse osmosis. Fine-tuning is spending a week with that librarian, showing them real questions and the right books. Hard negatives are the tricky cases: books that look relevant but are not. After the week, they shelve and find your books far better.
Why and when
Generic embedding models learn from broad web text. They may not understand your product names, jargon, abbreviations, or code-mixed queries like Roman Urdu. Embedding models are small and train quickly, often in minutes to an hour on one GPU, so fine-tuning them is cheap. But try the cheaper steps first: hybrid keyword and vector search, better chunking, metadata filters, and a reranker that rescores the top results. Fine-tune when retrieval still misses.
Contrastive training
How does it learn? You provide pairs: an anchor, which is a query, and a positive, the passage that answers it. The usual loss, multiple negatives ranking loss, pulls each query toward its own passage and pushes it away from every other passage in the batch. So bigger batches mean more negatives. And hard negatives, passages that look relevant but are not, teach the model the fine distinctions that matter in your domain.
Pair sources
Where do pairs come from? Search logs, where a query led to a document the user clicked and stayed on. Support tickets, where an agent linked a help article to a customer's question. And synthetic questions, where a model writes realistic questions for each passage, which you then filter. Real pairs from your own users are the most valuable, because they capture how people actually ask.
Hands-on: train_embed.py
The script in the lesson uses Sentence Transformers. It loads a multilingual base model, reads anchor and positive pairs, holds out ten percent, and optionally mines hard negatives with the current model. It builds an information retrieval evaluator, measures before training, trains one epoch with multiple negatives ranking loss and a no-duplicates batch sampler, then measures again and saves. Recent versions reorganized some import paths, so check the docs if an import fails.
Operations
Now the operational steps people forget. Re-index everything. Vectors from the old and new model are not comparable, so re-embed the whole corpus, ideally into a new index, then switch. Version the index with the model name and revision. Keep any query and document prefixes consistent between training and serving. And note that embedding models are small enough for CPU at moderate traffic, or you can serve them from inference servers that support embeddings.
Evaluate properly
Evaluate properly. The quick evaluator in the script uses only the test pairs, which is a start. A realistic evaluation searches the full corpus, all your passages, with real held-out queries, and reports recall at k and mean reciprocal rank, compared with the base model and with a reranker-only baseline. Also check that other languages did not regress.
Worked example: Roman Urdu search (illustrative)
A worked example, with illustrative numbers. A Pakistani fintech's help center is written in English, but many searches are in Roman Urdu, like bill payment nahi ho rahi. A generic multilingual model found the right article in the top five for about half of those. They built eight thousand pairs from agent-linked articles, added four thousand filtered synthetic questions, mined hard negatives, and trained for under an hour on one GPU. Recall at five on real Roman Urdu queries rose substantially, English held steady, and they switched to a new versioned index after a shadow comparison.
Simple example: electronics search
A simple example. An electronics store's search struggles when customers type phone charger that does not overheat, because product pages talk about thermal protection. They collect three hundred pairs from search logs where customers clicked and bought, fine-tune a small multilingual embedding model for twenty minutes, and re-index their two thousand product pages. Recall at five on real held-out searches improves noticeably, and they switch after a week of shadow testing.
Reranker or fine-tune?
Should you fine-tune the embedding model or add a reranker? Often both, in that order of cheapness. A reranker, a cross-encoder that reads the query and each candidate together, can fix ordering problems quickly without re-indexing. But a reranker can only reorder what retrieval already found. If the right passage never appears in the top fifty, only a better embedding model, or hybrid keyword search, will bring it in. Measure recall at fifty to tell which problem you have.
FAQ: tune embeddings, generator or both?
A question from RAG teams: should we fine-tune the embedding model, the generator, or both? Start with whichever step is failing. If the right passages are not being retrieved, the generator cannot fix that, so improve retrieval first. If the right passages are retrieved but answers are still poor, work on the generator's prompt, format or, eventually, fine-tuning. Your evaluation should separately measure retrieval hit rate and answer quality so you know which one to fix.
Try this now
Try this now. Export fifty real queries from your search or support logs, and for each, write the document that should have been found. Run them through your current retrieval and count how often the right document is in the top five. That single number is your baseline, and it tells you whether embedding fine-tuning is worth it.
Watch me do it (illustrative)
Watch me do it. I export three months of help-center searches where users clicked an article and did not come back to search again. After cleaning, I have two thousand four hundred query and article pairs. I split off ten percent, and for evaluation I use the full corpus of all articles, not just the held-out positives. Baseline recall at five: fifty-eight percent. I mine one hard negative per pair with the current model, and train one epoch with multiple negatives ranking loss and the no-duplicates sampler, batch size sixty-four. It takes twenty-five minutes. Recall at five on the same evaluation rises to seventy-three percent, and English-only queries do not regress. Then operations: I embed every article with the new model into a new index named with the model version, run both indexes in shadow for a week, compare results on live queries, and switch. All numbers illustrative.
Recap
Recap. Retrieval caps RAG quality. Try cheaper fixes first, then fine-tune embeddings with anchor and positive pairs, in-batch negatives and hard negatives. Evaluate on real queries against the full corpus, and always re-index into a versioned index. Your next step: build five hundred pairs from your logs, fine-tune, and report recall at five and MRR before and after.
Key takeaways
- Retrieval quality caps RAG quality; try hybrid search and rerankers first
- Contrastive training with (anchor, positive) pairs and MNRL uses in-batch negatives
- Hard negatives teach fine distinctions; avoid false negatives from duplicates
- Evaluate Recall@k/MRR on real queries against the full corpus
- Re-embed the whole corpus into a new, versioned index after fine-tuning
Try it
Build 500 (query, passage) pairs from your search or support logs, fine-tune an embedding model, and report Recall@5 and MRR before and after on real held-out queries.