Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Multimodal inputs and embeddings · lesson 8 of 19 · 15 min
Embeddings APIs and semantic search in production
What embeddings APIs give you
An embedding is a vector of numbers representing the meaning of text (or images). Similar meanings produce nearby vectors, which powers semantic search, clustering, deduplication, recommendations and retrieval for RAG.
Where to get embeddings (verify current model names and dimensions in docs):
- OpenAI:
client.embeddings.create(model="text-embedding-3-small" | "text-embedding-3-large", input=[...]), with an optionaldimensionsparameter to shorten vectors. - Google Gemini:
client.models.embed_content(model="gemini-embedding-001", contents=[...], config=EmbedContentConfig(output_dimensionality=...)), with task types for retrieval queries versus documents. - Anthropic does not offer its own embeddings model; its documentation points to third-party providers such as Voyage AI.
- Open-weight models (for example multilingual sentence-embedding models on Hugging Face) via hosted inference or self-hosting.
- Cloud platforms (Bedrock, Microsoft Foundry, Google's enterprise platform) host several embedding models.
Choosing an embedding model
- Language coverage: test Arabic, Urdu and English queries against mixed-language documents if that's your reality.
- Domain fit: product catalogs, legal text, code and chat behave differently.
- Dimensions: larger vectors can be more accurate but cost more storage and search time; many models let you reduce dimensions.
- Price and rate limits: embedding a million documents is a batch job (see lesson 9).
- Lock-in: vectors from different models are not compatible. Changing models means re-embedding everything. Store the model name and version with every vector.
Production pipeline
- Chunk documents (by headings or ~200–800 tokens with overlap), keeping metadata (source, date, language, permissions).
- Embed chunks in batches; retry failures; record model and dimensions.
- Store in a vector index: a Postgres extension such as pgvector, a search engine with vector support, or a dedicated vector database.
- Query: embed the query with the same model (using the query task type where supported), search with metadata filters, and combine with keyword search (hybrid) for names, codes and SKUs.
- Rerank top results with a reranker model when precision matters.
- Evaluate retrieval recall on a labeled question set before tuning generation.
Hands-on: semantic search over FAQs with two providers
import os, numpy as np
from openai import OpenAI
from google import genai
from google.genai import types
FAQS = ["How do I pause my subscription?", "Can I change my delivery address?",
"What payment methods do you accept?", "How do I get a refund for a damaged item?"]
def cosine(a, b):
a, b = np.array(a), np.array(b)
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
# OpenAI embeddings
oai = OpenAI()
E_MODEL = os.environ.get("OPENAI_EMBED_MODEL", "text-embedding-3-small")
doc_vecs = [d.embedding for d in oai.embeddings.create(model=E_MODEL, input=FAQS).data]
q_vec = oai.embeddings.create(model=E_MODEL, input=["my parcel arrived broken, money back?"]).data[0].embedding
print("OpenAI best:", max(zip(FAQS, doc_vecs), key=lambda x: cosine(q_vec, x[1]))[0])
# Gemini embeddings (document vs query task types)
g = genai.Client()
G_MODEL = os.environ.get("GEMINI_EMBED_MODEL", "gemini-embedding-001")
docs = g.models.embed_content(model=G_MODEL, contents=FAQS,
config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"))
q = g.models.embed_content(model=G_MODEL, contents="my parcel arrived broken, money back?",
config=types.EmbedContentConfig(task_type="RETRIEVAL_QUERY"))
g_docs = [e.values for e in docs.embeddings]
print("Gemini best:", max(zip(FAQS, g_docs), key=lambda x: cosine(q.embeddings[0].values, x[1]))[0])
Both should match the refund FAQ despite sharing almost no words with it. For production, store vectors in an index rather than recomputing, and check the docs for batch limits and current model names.
Worked example: a multilingual help center in the Gulf
A retailer serving the UAE and KSA has 1,200 help articles in Arabic and English. Customers ask in either language, often mixing them. The team evaluated two embedding models on 150 real queries (labeled with the correct article), chose the one with better cross-language recall, stored model name with every vector, and added keyword search for order numbers and product codes. Recall on the test set improved over keyword-only search, and "no result" searches dropped (illustrative).
Keeping the index fresh and safe
- Re-embedding triggers: document edits, deletions and permission changes should update or remove vectors promptly; stale vectors return outdated or forbidden content.
- Deletion requests: when a customer asks for their data to be deleted, remember derived data such as embeddings of their messages.
- Versioning: when you change chunking or models, build a new index alongside the old one, evaluate, then switch traffic; keep the old one until you are confident.
- Monitoring: log queries with zero or low-similarity results; they reveal missing content and new customer vocabulary.
Pitfalls
- Mixing vectors from different models or versions in one index.
- Embedding queries with a different model than documents.
- Pure semantic search for exact identifiers (use hybrid).
- Forgetting permission filters, so users retrieve documents they shouldn't see.
Measuring success
Recall@k and MRR on a labeled query set per language, zero-result rate, latency per query, and embedding cost per thousand documents.
Video lecture: Embeddings APIs and semantic search in production
Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.
- Embeddings and semantic search
- Why it matters
- The concept
- Where to get embeddings
- Simple example: four FAQs
- Choosing a model
- Production pipeline
- Keep the index healthy
- Business example: Gulf help center
- Measuring search
- Common mistakes
- Read your failed searches
- Deeper: the Gulf help center (illustrative)
- Watch me do it: FAQ semantic search
- Recap + try this now
Lecture transcript
Embeddings and semantic search
Your customer types, my parcel arrived broken, can I get my money back? Your help article is titled, refunds for damaged items. They share almost no words, so keyword search misses it. Embeddings fix that by matching meaning instead of words. In this lesson, you'll learn how to use embeddings APIs from different providers, how to build semantic search that holds up in production, and how to choose the right model for your languages.
Why it matters
Why does this matter? Search quality quietly drives customer satisfaction, support costs and the quality of every retrieval augmented AI answer. Think of a librarian who's read every book versus a card catalog that only knows titles. Ask the catalog about ocean pollution and it finds books with those exact words. Ask the librarian and they bring you a book called The Plastic Tide. Embeddings give your search the librarian's understanding.
The concept
Here's the concept. An embedding model turns a piece of text into a list of numbers, a vector, positioned so that similar meanings sit close together. To search, you embed every document once, embed each query when it arrives, and find the documents whose vectors are nearest to the query's vector, usually with cosine similarity. The model does the understanding; the vector index does the fast lookup.
Where to get embeddings
Where do you get embeddings? OpenAI offers text embedding three small and large, with an option to shorten vectors. Google's Gemini API offers a Gemini embedding model, with task types for queries versus documents. Anthropic doesn't offer its own embeddings model, and its documentation points to partners such as Voyage AI. You can also use open weight multilingual models, or embedding models hosted on the major cloud platforms. Model names change, so check the docs before you commit.
Simple example: four FAQs
A simple example from the lesson code. Four FAQs: pausing a subscription, changing an address, payment methods and refunds for damaged items. The query: my parcel arrived broken, money back? The code embeds the FAQs and the query with OpenAI, then again with Gemini using document and query task types, and picks the most similar FAQ. Both find the refund article, even though it shares almost no words with the question.
Choosing a model
How do you choose a model? Test language coverage with your real mix, like Arabic, Urdu and English, including mixed language queries. Check domain fit: product catalogs, legal text and chat behave differently. Consider dimensions, since bigger vectors can be more accurate but cost storage and search time. Plan for price and rate limits, because embedding a million documents is a batch job. And remember lock in: vectors from different models aren't compatible, so switching models means re embedding everything.
Production pipeline
Now the production pipeline. Chunk documents by headings or a few hundred tokens with overlap, keeping metadata like source, date, language and permissions. Embed chunks in batches and record the model name and dimensions with every vector. Store them in an index, such as pgvector in Postgres or a vector database. At query time, embed with the same model, filter by metadata, and combine with keyword search for names, codes and product numbers. Rerank the top results when precision matters.
Keep the index healthy
Let's talk about keeping the index healthy, because search quality decays quietly. When a help article is edited, its old vectors still point to the old wording. When a document is deleted, or its permissions change, its vectors must go too, or users will find content they shouldn't. When a customer asks for their data to be deleted, remember embeddings of their messages count as their data. And when you change the model or chunking, build a new index alongside the old one, evaluate it, then switch.
Business example: Gulf help center
Now a realistic business example, with illustrative results. A retailer serving the UAE and Saudi Arabia has twelve hundred help articles in Arabic and English, and customers often mix both languages in one query. The team labeled a hundred and fifty real queries with the correct article, compared two embedding models, and chose the one with better cross language recall. They stored the model name with every vector and added keyword search for order numbers and product codes. Recall improved over keyword only search, and searches with no result dropped.
Measuring search
How do you measure it? Build a labeled set of real queries, each with the correct document. Measure recall at three or five: how often the right document appears in the top results. Measure mean reciprocal rank to see how high it appears. Track the zero result rate, latency per query and embedding cost per thousand documents. Measure all of this per language, because averages hide the customers you're failing.
Common mistakes
Common mistakes. Mixing vectors from different models or versions in one index. Embedding queries with a different model than the documents. Relying purely on semantic search for exact identifiers. And forgetting permission filters, so users retrieve documents they shouldn't see. The last one is a security issue, not just a quality issue.
Read your failed searches
One more habit worth building: read your failed searches. Log every query that returns nothing, or only weak matches, and review a sample every week. You'll discover missing articles, new product names customers use, and phrases in Urdu or Arabic your content never mentions. Each finding is either a new article, a synonym, or a new evaluation case. Over a few months, this simple loop does more for search quality than swapping embedding models.
Deeper: the Gulf help center (illustrative)
Let's deepen the Gulf help center example with illustrative numbers. A hundred and fifty labeled queries, roughly half Arabic, a third English, and the rest mixed. Keyword search found the right article in the top three for about half of queries. The first embedding model tested lifted English recall well but did poorly on mixed queries. The second model performed well across all three groups, so they chose it, even though it cost a little more per thousand documents. Adding keyword search back for order numbers, in a hybrid setup, fixed the last common failure. Zero result searches fell sharply, and contact center tickets tagged, couldn't find the answer, dropped the following month.
Watch me do it: FAQ semantic search
Watch me do it. Let's walk through the FAQ search script. Four FAQs sit in a list. The cosine function turns two lists into arrays and returns their dot product divided by the product of their lengths. For OpenAI, I embed all four FAQs in one call, collect the vectors, then embed the customer query, my parcel arrived broken, money back. I pick the FAQ with the highest cosine similarity and print it: the refund FAQ. For Gemini, I call embed content with the FAQs and a task type of retrieval document, then embed the query with retrieval query, because Gemini tunes vectors differently for documents and questions. I compare again, and the refund FAQ wins again. Notice two rules in action: queries and documents use the same model, and I never compare an OpenAI vector with a Gemini vector. In production, the document vectors would be stored in an index with the model name, not recomputed each time.
Recap + try this now
Quick recap. Embeddings match meaning. OpenAI and Gemini offer embeddings APIs, Anthropic points to partners, and open models and cloud platforms are options too. Store the model with every vector, use hybrid search and permission filters, and measure recall per language. Try this now: embed fifty of your own FAQs with two models, write thirty labeled test queries in your customers' languages, and compare recall at three.
Key takeaways
- Embeddings map meaning to vectors for search, clustering, deduplication and RAG retrieval.
- OpenAI and Gemini offer embeddings APIs; Anthropic points to partners such as Voyage AI; open models and cloud platforms are options too.
- Vectors from different models are incompatible: store model name and version, and re-embed when switching.
- Use hybrid search for exact identifiers and permission filters for security.
- Evaluate recall per language on labeled queries before tuning generation.
Try it
Embed 50 of your FAQs or help articles with two embedding models, write 30 labeled test queries (in your customers' languages), and compare recall@3.