---
title: "Embeddings and semantic search | Optimize All Academy"
description: "The core idea An embedding is a list of numbers (a vector) that represents the meaning of a piece of text, an image or other data. An embedding model is…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/embeddings-and-semantic-search
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Embeddings, semantic search and indexing · lesson 1 of 20 · 12 min

# Embeddings and semantic search

## The core idea

An **embedding** is a list of numbers (a vector) that represents the meaning of a piece of text, an image or other data. An embedding model is trained so that items with similar meaning end up close together in this vector space. "How do I get a refund?" and "Can I return my order for my money back?" share few words but land near each other; "refund" and "refinery" share letters but land far apart.

This enables **semantic search**: instead of matching keywords, you embed the query, then find the stored items whose vectors are closest to it.

## How similarity is measured

The usual measure is **cosine similarity**: the cosine of the angle between two vectors, where closer to 1 means more similar in direction. Many systems normalise vectors so cosine similarity and dot product give the same ranking. You rarely compute this yourself; vector databases and libraries handle it. What matters is understanding that "nearest" means "most similar according to this embedding model", which may not be what your users mean by relevant.

```python
import numpy as np

def cosine(a, b):
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

# query_vec and doc_vecs come from the same embedding model
scores = [cosine(query_vec, d) for d in doc_vecs]
top_k = np.argsort(scores)[::-1][:5]
```

## Vector databases and approximate search

With thousands of items, you can compare the query against everything. With millions, systems use **approximate nearest neighbour (ANN)** indexes that trade a tiny amount of accuracy for large speed gains. Options range from extensions to databases you may already run, to dedicated vector databases, to in-memory libraries. Choose based on scale, operational comfort, filtering needs and cost, not hype. For many business knowledge bases, a vector extension on an existing database is enough.

## Semantic vs keyword vs hybrid search

- **Keyword search** (for example BM25-style ranking) excels at exact terms: product codes, names, error messages, legal clause numbers.
- **Semantic search** excels at paraphrases, questions phrased differently from the documents, and cross-lingual matching with multilingual models.
- **Hybrid search** runs both and merges results (for example with reciprocal rank fusion). In practice, hybrid is a strong default for business content, because users ask both "What's our policy on late delivery?" and "What is SKU 44-A?"

## Choosing an embedding model

Consider:

- **Language coverage.** If your content or users span English, Arabic and Urdu, test a multilingual model on real queries in each.
- **Domain fit.** General models may underperform on specialised jargon. Test before assuming.
- **Dimensions and cost.** Larger vectors can capture more nuance but cost more to store and search. Some models support shortened vectors with modest quality loss.
- **Stability.** If you change embedding models, you must re-embed your entire corpus; vectors from different models are not comparable.

Public leaderboards are a starting point, but your own queries on your own documents are the only benchmark that matters.

## Metadata and filtering

Store metadata alongside each vector: source, date, author, product, region, access level. Filtering by metadata (for example "only documents for the UAE office, updated this year, that this user may see") often improves relevance more than a better embedding model, and it is essential for permissions.

## Worked example: a support knowledge base

A software company has 3,000 help articles. Keyword search fails on questions like "the app keeps logging me out", because the article is titled "Session timeout settings". With semantic search, the question lands near the right article. But searching "error E1042" semantically returns vaguely related articles about errors in general. Hybrid search with metadata filters (product version) handles both.

## Failure modes

- **Semantic near-misses.** "How to cancel my subscription" retrieves "How to subscribe" because the topics are close. Rerankers (next module) help.
- **Negation blindness.** Embeddings often poorly capture "not" or "except".
- **Stale vectors.** Documents update but embeddings do not. Build re-indexing into your content pipeline.
- **Mixed-model vectors.** Embedding queries with a different model than the documents silently breaks retrieval.

## Hands-on: build hybrid search in 40 lines

This runs on a laptop with no API key, using open-source models. It shows keyword search, semantic search and reciprocal rank fusion (RRF) side by side, so you can see each one win and lose.

```bash
python -m venv .venv && source .venv/bin/activate
pip install sentence-transformers rank-bm25 numpy
```

```python
import numpy as np
from sentence_transformers import SentenceTransformer
from rank_bm25 import BM25Okapi

docs = [
    {"id": "kb-01", "text": "Session timeout settings: the app signs you out after 30 minutes of inactivity."},
    {"id": "kb-02", "text": "Error E1042 means your payment card was declined by the bank."},
    {"id": "kb-03", "text": "Refunds: return unused items within 14 days for a full refund."},
    {"id": "kb-04", "text": "سياسة الاسترداد: يمكنك إرجاع المنتجات غير المستخدمة خلال 14 يومًا."},
]

# A small multilingual model; swap in others and compare on YOUR queries.
model = SentenceTransformer("sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2")
doc_vecs = model.encode([d["text"] for d in docs], normalize_embeddings=True)
bm25 = BM25Okapi([d["text"].lower().split() for d in docs])

def semantic(query, k=3):
    q = model.encode([query], normalize_embeddings=True)[0]
    scores = doc_vecs @ q  # dot product = cosine similarity for normalised vectors
    return [docs[i]["id"] for i in np.argsort(-scores)[:k]]

def keyword(query, k=3):
    scores = bm25.get_scores(query.lower().split())
    return [docs[i]["id"] for i in np.argsort(-scores)[:k]]

def rrf(*rankings, k=60):
    """Reciprocal rank fusion: reward documents that rank well in any list."""
    fused = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
    return sorted(fused, key=fused.get, reverse=True)

for q in ["the app keeps logging me out", "E1042", "can I get my money back?"]:
    print(f"{q!r:35} semantic={semantic(q)} keyword={keyword(q)} hybrid={rrf(semantic(q), keyword(q))[:3]}")
```

What to notice when you run it:

- "the app keeps logging me out" shares no useful words with kb-01, so keyword search struggles while semantic search finds it.
- "E1042" is an exact code: keyword search nails it; semantic search may rank it lower.
- "can I get my money back?" should surface both the English and the Arabic refund articles semantically, because the model is multilingual.
- RRF needs no score calibration between the two systems, which is why it is such a common default.

## Measure before you choose

Turn the activity at the end of this lesson into a tiny benchmark. Label 20 to 50 real questions with the document that answers them, then compute **recall at 3** for keyword, semantic and hybrid:

```python
labelled = [("the app keeps logging me out", "kb-01"), ("E1042", "kb-02"), ("can I get my money back?", "kb-03")]
def recall_at_k(search, k=3):
    return sum(gold in search(q)[:k] for q, gold in labelled) / len(labelled)
print("keyword", recall_at_k(keyword), "semantic", recall_at_k(semantic),
      "hybrid", recall_at_k(lambda q: rrf(semantic(q), keyword(q))))
```

In production you would move the vectors into a store such as the pgvector extension for PostgreSQL, a managed vector database or your search engine's vector features, but this measurement habit stays the same.

## Going further

Embeddings power more than search: clustering feedback into themes, deduplicating content, recommending similar items, routing tickets and selecting few-shot examples dynamically. The same caveats apply: validate on your data and keep humans reviewing high-impact uses.

## Video lecture: Embeddings and semantic search

Lecture coming soon · 15 chapters · about 10 minutes. Read the full transcript below.

1. Why search misses the obvious
2. Why it matters
3. Embeddings = meaning as coordinates
4. Measuring similarity
5. Simple example
6. Keyword vs semantic vs hybrid
7. Choosing an embedding model
8. Metadata filters matter
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. How you'll know it's working
13. Watch me do it: hybrid search script
14. Recap
15. Try this now (20 minutes)

## Lecture transcript

### Why search misses the obvious

Here's a puzzle. A customer types, the app keeps logging me out. Your help centre has the perfect article, titled Session timeout settings. Old-fashioned keyword search returns nothing useful, because the two sentences share no important words. By the end of this lesson you'll understand exactly why that happens, how embeddings fix it, where embeddings fail, and how to combine both approaches so your search, and later your AI assistant, finds the right answer far more often.

### Why it matters

Why should you care? Because almost every AI feature you'll build in this course starts with finding the right information. A support assistant, a sales research agent, a policy chatbot: if search fails, everything downstream fails, however clever the model. Here's an analogy. Keyword search is like a librarian who only matches the exact words on the spine. Semantic search is like a librarian who understands what you mean, even if you describe the book badly. The best libraries employ both.

### Embeddings = meaning as coordinates

An embedding is a list of numbers that represents meaning. An embedding model has been trained on huge amounts of text so that things which mean similar things end up close together in a space with hundreds or thousands of dimensions. You can't picture that many dimensions, so imagine a map instead. How do I get a refund, and can I send my order back for my money, land in the same neighbourhood. Refund and refinery share letters, but they live in different cities. Search then becomes geography: turn the question into a point on the map, and look for the nearest documents.

### Measuring similarity

How do we measure near? The usual answer is cosine similarity, which compares the direction of two vectors. If you normalise your vectors, a plain dot product gives the same ranking, which is why many libraries do exactly that. At small scale you can compare the question with every document. At millions of documents, systems use approximate nearest neighbour indexes, such as HNSW graphs, trading a sliver of accuracy for a big jump in speed. For most business knowledge bases, a vector extension on a database you already run, like pgvector on Postgres, is plenty.

### Simple example

Let's make it concrete with a simple example. Take three sentences. One: my card was declined. Two: the payment didn't go through. Three: the card game was fun. An embedding model places one and two close together, because they mean nearly the same thing, and places three far away, even though it shares the word card. Now ask, why did my payment fail? The nearest neighbours are one and two. That's semantic search in a single picture, and it's exactly why it rescues paraphrased questions.

### Keyword vs semantic vs hybrid

Now the catch. Semantic search is brilliant at paraphrases and natural questions, but weak at exact things. Ask for error E1042 and it may return general articles about errors, because they all feel similar. Keyword search, the BM25 family of ranking, is the opposite: great for product codes, names, clause numbers and error messages, useless for paraphrases. That's why hybrid search is the strong default for business content. You run both searches and fuse the ranked lists. A popular method is reciprocal rank fusion, which rewards documents that rank highly in either list and needs no fiddly score calibration.

### Choosing an embedding model

Choosing an embedding model is where teams waste the most time reading leaderboards. Leaderboards are a starting point, not a verdict. What matters is how a model performs on your questions, in your users' languages, with your jargon. If your customers write in English, Arabic, Urdu and Roman Urdu, test all four. Remember three practical rules. Bigger vectors cost more to store and search. Some models let you shorten vectors with modest quality loss. And if you ever switch models, you must re-embed everything, because vectors from different models are not comparable.

### Metadata filters matter

There's one more lever that often beats a fancier model: metadata. Store the source, date, product, region and access level with every chunk, and filter on it. A Dubai employee asking about leave should only see the Dubai policy, and only documents they're allowed to open. Filtering improves relevance and enforces permissions at the same time. Without it, your AI assistant can become a very polite way for people to read documents they should never have seen.

### Business example (illustrative)

Let's deepen the business picture with an illustrative case. A Riyadh electronics retailer has around two thousand help articles in Arabic and English. Before hybrid search, its logs showed roughly one search in four ending with no click. After adding semantic search fused with keyword search, plus a filter for the customer's country, that dropped to about one in ten, and searches for exact model numbers still worked. The numbers are illustrative, but the pattern is typical: hybrid fixes paraphrases without breaking codes.

### Hands-on in the lesson

In the lesson text you'll find a hands-on script that runs on a laptop with no API key. It builds keyword search, semantic search and fused hybrid search over a handful of help articles, including one in Arabic, then measures recall at three, which simply asks: for what share of questions did the right document appear in the top three results? Run it, then replace the sample articles with twenty of your own and twenty real questions. That tiny benchmark will teach you more about your data than any vendor demo.

### Common mistakes

Before we recap, the most common mistakes. First, testing on questions you wrote yourself, which mirror your documents' wording and make any search look good. Second, switching the query embedding model without re-embedding the documents. Third, forgetting metadata filters, so users retrieve documents from the wrong region or ones they shouldn't see. Fourth, ignoring languages: an English-only model quietly failing your Arabic or Urdu speakers. And fifth, choosing a vector database before measuring anything. Measure first, then choose infrastructure.

### How you'll know it's working

How will you know your search is working? Track three numbers. Recall at three or five on your labelled questions, which should rise as you move from keyword to hybrid. The share of real user searches that end in a click or a solved question. And the no-results rate, which should fall. If recall rises but users still can't find things, your test questions don't reflect real behaviour, so add more real queries from your logs.

### Watch me do it: hybrid search script

Watch me do it. I open the script from the lesson. First, the four documents, one of them in Arabic. Next, I load the multilingual model and call encode with normalise embeddings set to true, so a dot product equals cosine similarity. Then I build the BM25 index by splitting each document into lower-case words. The semantic function embeds the query and sorts documents by score. The keyword function asks BM25 for scores. And rrf adds one over sixty plus rank for every list a document appears in. I run it. For the app keeps logging me out, keyword search misses, semantic search finds kb-01, and hybrid keeps it on top. For E1042, keyword wins and hybrid still ranks kb-02 first. For can I get my money back, both refund articles appear, including the Arabic one. Finally, I run recall at k on the three labelled questions and see hybrid score highest. That's the whole loop: build, run, measure.

### Recap

Let's recap. Embeddings turn meaning into coordinates, and semantic search finds the nearest ones. Keyword search wins on exact terms, semantic wins on paraphrases and languages, and hybrid with rank fusion gives you both. Choose models by testing on your own questions, store and filter on metadata, and re-embed everything if you change models. Your next step: collect twenty real questions from your inbox or help desk, label the right answer for each, and measure. Next, we'll look at chunking, the quiet decision that controls what your system can find in the first place.

### Try this now (20 minutes)

Try this now. Open your support inbox, help desk or sales enquiries and copy ten real questions. For each one, note which article or document should answer it. Then run the hands-on script, replacing the sample articles with those documents, and record how many questions get the right answer in the top three for keyword, semantic and hybrid search. It takes about twenty minutes, and you'll have your first retrieval benchmark, which you'll reuse in every lesson that follows.

## Video transcript

In this course we're going to build up, piece by piece, the techniques behind today's most capable AI systems: retrieval, tools, agents and the Model Context Protocol. And it all starts with one idea: embeddings.

An embedding is a list of numbers that represents meaning. An embedding model is trained so that things with similar meaning end up close together. So "How do I get a refund?" and "Can I send my order back for my money?" land near each other, even though they share almost no words.

That gives us semantic search. You embed the question, then look for the stored passages whose vectors are nearest. It's brilliant for paraphrases and natural questions. But it has blind spots. It can miss exact things like product codes or error numbers, and it's often weak at negation. That's why, for most business content, the strong default is hybrid search: keyword and semantic together, merged into one ranked list.

Three practical rules to remember. First, choose your embedding model by testing it on your own questions, in your users' languages, not by reading a leaderboard. Second, store metadata with every chunk, such as source, date, region and who's allowed to see it, and filter on it. That often improves relevance more than a fancier model, and it's how you enforce permissions. Third, if you change embedding models, you re-embed everything. Vectors from different models don't mix.

Next, we'll look at chunking: how you cut your documents up, which quietly decides what your AI can and can't find.

## Key takeaways

- Embeddings map meaning to vectors; semantic search finds nearest vectors to the query.
- Keyword search wins on exact terms, semantic on paraphrases; hybrid is a strong default for business content.
- Choose embedding models by testing on your own queries, languages and domain.
- Metadata filtering improves relevance and enforces permissions; re-embed everything if you change models.

## Try it

Collect 20 real questions and the article that answers each. Compare keyword-only and semantic (or hybrid) search: how many correct articles appear in the top 3?

- [Next: Chunking and indexing strategies](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/chunking-and-indexing)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
