Latest AI Techniques: RAG, Tool Use, Agents & MCPEmbeddings, semantic search and indexing · Lesson 1 of 20

Embeddings and semantic search

Video lesson · 12 min · 10 min lecture

Video lecture

Embeddings and semantic search

15 chapters · about 10 min · full transcript

Coming soon

Chapter 1 of 15

Why search misses the obvious

  • Keyword match vs meaning match
  • Embeddings, similarity, hybrid search
  • How to measure which one wins

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The core idea

An embedding is a list of numbers (a vector) that represents the meaning of a piece of text, an image or other data. An embedding model is trained so that items with similar meaning end up close together in this vector space. "How do I get a refund?" and "Can I return my order for my money back?" share few words but land near each other; "refund" and "refinery" share letters but land far apart.

This enables semantic search: instead of matching keywords, you embed the query, then find the stored items whose vectors are closest to it.

How similarity is measured

The usual measure is cosine similarity: the cosine of the angle between two vectors, where closer to 1 means more similar in direction. Many systems normalise vectors so cosine similarity and dot product give the same ranking. You rarely compute this yourself; vector databases and libraries handle it. What matters is understanding that "nearest" means "most similar according to this embedding model", which may not be what your users mean by relevant.

import numpy as np

def cosine(a, b):
    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

# query_vec and doc_vecs come from the same embedding model
scores = [cosine(query_vec, d) for d in doc_vecs]
top_k = np.argsort(scores)[::-1][:5]

With thousands of items, you can compare the query against everything. With millions, systems use approximate nearest neighbour (ANN) indexes that trade a tiny amount of accuracy for large speed gains. Options range from extensions to databases you may already run, to dedicated vector databases, to in-memory libraries. Choose based on scale, operational comfort, filtering needs and cost, not hype. For many business knowledge bases, a vector extension on an existing database is enough.

  • Keyword search (for example BM25-style ranking) excels at exact terms: product codes, names, error messages, legal clause numbers.
  • Semantic search excels at paraphrases, questions phrased differently from the documents, and cross-lingual matching with multilingual models.
  • Hybrid search runs both and merges results (for example with reciprocal rank fusion). In practice, hybrid is a strong default for business content, because users ask both "What's our policy on late delivery?" and "What is SKU 44-A?"

Choosing an embedding model

Consider:

  • Language coverage. If your content or users span English, Arabic and Urdu, test a multilingual model on real queries in each.
  • Domain fit. General models may underperform on specialised jargon. Test before assuming.
  • Dimensions and cost. Larger vectors can capture more nuance but cost more to store and search. Some models support shortened vectors with modest quality loss.
  • Stability. If you change embedding models, you must re-embed your entire corpus; vectors from different models are not comparable.

Public leaderboards are a starting point, but your own queries on your own documents are the only benchmark that matters.

Metadata and filtering

Store metadata alongside each vector: source, date, author, product, region, access level. Filtering by metadata (for example "only documents for the UAE office, updated this year, that this user may see") often improves relevance more than a better embedding model, and it is essential for permissions.

Worked example: a support knowledge base

A software company has 3,000 help articles. Keyword search fails on questions like "the app keeps logging me out", because the article is titled "Session timeout settings". With semantic search, the question lands near the right article. But searching "error E1042" semantically returns vaguely related articles about errors in general. Hybrid search with metadata filters (product version) handles both.

Failure modes

  • Semantic near-misses. "How to cancel my subscription" retrieves "How to subscribe" because the topics are close. Rerankers (next module) help.
  • Negation blindness. Embeddings often poorly capture "not" or "except".
  • Stale vectors. Documents update but embeddings do not. Build re-indexing into your content pipeline.
  • Mixed-model vectors. Embedding queries with a different model than the documents silently breaks retrieval.

Hands-on: build hybrid search in 40 lines

This runs on a laptop with no API key, using open-source models. It shows keyword search, semantic search and reciprocal rank fusion (RRF) side by side, so you can see each one win and lose.

python -m venv .venv && source .venv/bin/activate
pip install sentence-transformers rank-bm25 numpy
import numpy as np
from sentence_transformers import SentenceTransformer
from rank_bm25 import BM25Okapi

docs = [
    {"id": "kb-01", "text": "Session timeout settings: the app signs you out after 30 minutes of inactivity."},
    {"id": "kb-02", "text": "Error E1042 means your payment card was declined by the bank."},
    {"id": "kb-03", "text": "Refunds: return unused items within 14 days for a full refund."},
    {"id": "kb-04", "text": "سياسة الاسترداد: يمكنك إرجاع المنتجات غير المستخدمة خلال 14 يومًا."},
]

# A small multilingual model; swap in others and compare on YOUR queries.
model = SentenceTransformer("sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2")
doc_vecs = model.encode([d["text"] for d in docs], normalize_embeddings=True)
bm25 = BM25Okapi([d["text"].lower().split() for d in docs])

def semantic(query, k=3):
    q = model.encode([query], normalize_embeddings=True)[0]
    scores = doc_vecs @ q  # dot product = cosine similarity for normalised vectors
    return [docs[i]["id"] for i in np.argsort(-scores)[:k]]

def keyword(query, k=3):
    scores = bm25.get_scores(query.lower().split())
    return [docs[i]["id"] for i in np.argsort(-scores)[:k]]

def rrf(*rankings, k=60):
    """Reciprocal rank fusion: reward documents that rank well in any list."""
    fused = {}
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking):
            fused[doc_id] = fused.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
    return sorted(fused, key=fused.get, reverse=True)

for q in ["the app keeps logging me out", "E1042", "can I get my money back?"]:
    print(f"{q!r:35} semantic={semantic(q)} keyword={keyword(q)} hybrid={rrf(semantic(q), keyword(q))[:3]}")

What to notice when you run it:

  • "the app keeps logging me out" shares no useful words with kb-01, so keyword search struggles while semantic search finds it.
  • "E1042" is an exact code: keyword search nails it; semantic search may rank it lower.
  • "can I get my money back?" should surface both the English and the Arabic refund articles semantically, because the model is multilingual.
  • RRF needs no score calibration between the two systems, which is why it is such a common default.

Measure before you choose

Turn the activity at the end of this lesson into a tiny benchmark. Label 20 to 50 real questions with the document that answers them, then compute recall at 3 for keyword, semantic and hybrid:

labelled = [("the app keeps logging me out", "kb-01"), ("E1042", "kb-02"), ("can I get my money back?", "kb-03")]
def recall_at_k(search, k=3):
    return sum(gold in search(q)[:k] for q, gold in labelled) / len(labelled)
print("keyword", recall_at_k(keyword), "semantic", recall_at_k(semantic),
      "hybrid", recall_at_k(lambda q: rrf(semantic(q), keyword(q))))

In production you would move the vectors into a store such as the pgvector extension for PostgreSQL, a managed vector database or your search engine's vector features, but this measurement habit stays the same.

Going further

Embeddings power more than search: clustering feedback into themes, deduplicating content, recommending similar items, routing tickets and selecting few-shot examples dynamically. The same caveats apply: validate on your data and keep humans reviewing high-impact uses.

Key takeaways

  • Embeddings map meaning to vectors; semantic search finds nearest vectors to the query.
  • Keyword search wins on exact terms, semantic on paraphrases; hybrid is a strong default for business content.
  • Choose embedding models by testing on your own queries, languages and domain.
  • Metadata filtering improves relevance and enforces permissions; re-embed everything if you change models.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A user searches 'error E1042' and semantic search returns general error articles. What most directly fixes this?
  2. You switch to a new embedding model for queries only. What happens?
  3. Why store metadata such as region and access level with each vector?

Put it into practice

Collect 20 real questions and the article that answers each. Compare keyword-only and semantic (or hybrid) search: how many correct articles appear in the top 3?

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.