Evaluating and Monitoring LLM ApplicationsMetrics and graders: code, judges and quality dimensions · Lesson 6 of 16
Groundedness, safety, privacy and fairness metrics
Video lecture
Groundedness, safety, privacy and fairness metrics
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Groundedness, safety and quality metrics
The failures that make headlines are rarely about accuracy alone. They are a chatbot inventing a refund policy, giving dangerous advice, leaking a customer's phone number, or treating users differently depending on their language. In this lecture you will learn to evaluate groundedness, safety, refusals, privacy, fairness and tone, each with the right dataset and grader.
0:24 Analogy: the fact-checking editor
An analogy for groundedness. Think of a journalist writing a story from interview notes. A good editor goes through the draft line by line and asks: where does this claim come from? If a sentence has no source in the notes, it is either cut or checked. Claim-level groundedness is that editor, working on every answer your system produces, asking for each statement: which retrieved passage supports this?
0:54 Groundedness
Start with groundedness, also called faithfulness. A response is grounded when every factual claim is supported by the context it was given, such as retrieved documents or tool outputs. Ungrounded claims are what people call hallucinations, and they are the top failure of retrieval and support systems. The standard method has three steps: split the answer into atomic claims, check each claim against the context, and score the share of claims that are supported.
1:26 Hands-on groundedness
Open-source libraries such as Ragas and DeepEval implement variations of this, and the lesson includes a compact Python version you can adapt. It returns the score and, more usefully, the list of unsupported claims, which tells you exactly what went wrong. Calibrate it against human labels like any judge. And decide a policy question up front: is the assistant allowed to use general knowledge beyond the provided sources, or must everything be grounded?
1:58 Safety layers
Next, safety. Moderation endpoints and open safety classifiers flag categories like hate, harassment, self-harm and violence. They are a useful first layer, but check their definitions and language coverage, especially for Arabic and Urdu, before you trust them. Add policy judges based on your own rules, like no investment advice or no medication dosages. And build adversarial datasets with jailbreak attempts and sensitive topics.
2:26 Refusals: both directions
Refusals need measuring in both directions. The harmful-request refusal rate: does the system refuse or safely redirect when it should? And the over-refusal rate: does it refuse perfectly benign requests, like how do I kill a Python process? Tightening safety prompts often raises both. If you only track one, you will silently make the product worse on the other.
2:52 Privacy
Privacy next. Combine code detectors, like patterns for emails, phone numbers, Pakistani CNIC or Emirates ID formats and card numbers with a checksum, with named-entity tools such as Microsoft Presidio. Add adversarial prompts that try to extract other users' data or your system prompt. And remember your logs and traces: personal data in telemetry is also a leak.
3:17 Fairness and tone
Fairness and consistency: run counterfactual tests, asking the same question while changing only a name, country or language, and compare outcomes. Does an assistant give different loan information when the user writes in Urdu rather than English? Differences are not always bias, but they must be explainable. For tone and brand, turn your brand guide into binary criteria, or compare against approved reference answers pairwise.
3:45 Case: UAE clinic assistant
Here is a complete example. A clinic group in the UAE runs an assistant for appointments and general wellness questions. Its safety suite has three hundred adversarial prompts: dosage requests, emergency symptoms and requests for diagnoses. A code check ensures emergency keywords always trigger the emergency contact response, and that is blocking at one hundred percent. A policy judge enforces no diagnosis or dosage. Over-refusal tests use benign questions. Groundedness covers clinic information, and PII detection runs on outputs and logs.
4:20 Pitfalls
Four pitfalls to avoid. Treating moderation scores as complete safety, when they miss domain-specific harms like unlicensed financial advice. Ignoring over-refusal, which frustrates users and quietly pushes them to other tools. Testing safety only in English for a multilingual product. And forgetting that telemetry is a privacy surface, so a perfectly safe answer can still leak through a log.
4:46 Example: two claims, one supported
A simple example of a groundedness check. The context says: standard delivery to Karachi takes three to five working days. The answer says: your order will arrive in three to five working days, and delivery is free for orders over five thousand rupees. The checker splits it into two claims. The first is supported by the context. The second is not, because nothing about free delivery was retrieved. Score: one out of two. And the unsupported claim is exactly the one that could cost you money if a customer holds you to it.
5:26 Scenario: the over-refusal surprise (illustrative)
Now a realistic scenario with illustrative numbers. A fintech in Abu Dhabi measured only its harmful-request refusal rate and was proud that it refused almost every risky prompt. Then they built an over-refusal set of one hundred benign questions, like how do I close my account or what does chargeback mean. The assistant refused a surprisingly large share of them. Customers had been quietly giving up and phoning the call center. After tuning, harmful refusals stayed high and benign refusals dropped sharply. Two metrics, one honest picture.
6:04 Deeper: a counterfactual pair set
One level deeper on counterfactual testing. Take twenty eligibility questions, and ask each twice: once written in English, once in Urdu, with identical facts. Compare the key points in each pair. If the Urdu answers systematically omit a document requirement, that is a difference you must explain or fix, whether it comes from retrieval, translation or the model.
6:29 Watch me do it: groundedness on one answer
Watch me do it. I run the groundedness function on one answer. The context is two policy snippets: standard delivery to Karachi takes three to five working days, and orders can be tracked on the order page. The answer: your order should arrive in three to five working days, you can track it on the order page, and delivery is free on orders over five thousand rupees. Step one, decompose. The judge returns three claims. Step two, verify each one against the context. Claim one, three to five working days: supported, and it quotes the evidence. Claim two, tracking on the order page: supported. Claim three, free delivery over five thousand rupees: not supported, with empty evidence. Step three, score: two out of three. The unsupported claim list tells me exactly what to investigate, and I check whether the free-delivery promotion exists at all. It ended last month, and the model remembered it from an old document still sitting in the index. So the fix is content, not the prompt: remove the old document and add the current promotion page. Then I add the same case to the over-refusal and PII suite review, and finally I spot-check the judge's claim splitting on ten more answers, because this metric is only as good as its decomposition.
8:02 Recap
Recap. Measure groundedness claim by claim. Layer safety with classifiers, policy judges and adversarial sets, and check language coverage. Track refusals in both directions. Detect PII in outputs and telemetry. Test fairness with counterfactuals, and turn your brand guide into criteria. Your next step: add an over-refusal set and a PII check to your eval suite this week, since those are the two most commonly missing.
Beyond "is it correct?"
Production LLM systems need evaluation on several dimensions that are easy to forget until something goes wrong publicly: groundedness (is it supported by the provided sources?), safety (toxicity, harmful advice, policy violations), privacy (PII leakage), refusal behavior (refusing when it should, not refusing when it should not), bias and fairness across user groups, and brand and tone. Each needs specific datasets and graders.
Groundedness and faithfulness
A response is grounded (or faithful) when every factual claim is supported by the context provided (retrieved documents, tool outputs). Ungrounded claims, often called hallucinations, are the top failure of RAG and support systems.
A common method, used by open-source libraries such as Ragas and DeepEval in various forms:
- Decompose the answer into atomic claims.
- Verify each claim against the context (supported / not supported).
- Score = supported claims / total claims.
# groundedness.py: claim-level groundedness with an LLM (sketch; add retries and logging)
import json, os
from anthropic import Anthropic
client = Anthropic(); MODEL = os.environ["JUDGE_MODEL"]
def ask_json(prompt: str) -> dict:
msg = client.messages.create(model=MODEL, max_tokens=800, temperature=0,
messages=[{"role": "user", "content": prompt}])
t = msg.content[0].text
return json.loads(t[t.find("{"): t.rfind("}") + 1])
def groundedness(answer: str, context: str) -> dict:
claims = ask_json("Split the ANSWER into atomic factual claims. Ignore greetings and "
"questions. Return {\"claims\": [..]}.\nANSWER:\n" + answer)["claims"]
results = []
for c in claims:
r = ask_json("Is the CLAIM fully supported by the CONTEXT? Answer only from the context. "
"Return {\"supported\": true|false, \"evidence\": \"<quote or empty>\"}.\n"
f"CONTEXT:\n{context}\n\nCLAIM:\n{c}")
results.append({"claim": c, **r})
score = sum(r["supported"] for r in results) / len(results) if results else 1.0
return {"score": score, "unsupported": [r["claim"] for r in results if not r["supported"]]}Calibrate this like any judge (previous lesson). Watch for claims that come from general knowledge: decide whether your product allows them.
Safety and toxicity
- Classifier-based: use moderation endpoints or open safety classifiers (for example the moderation APIs offered by major providers, or open models such as Llama Guard-style classifiers) to flag categories like hate, harassment, self-harm, sexual content and violence. Check category definitions and language coverage for Arabic and Urdu before trusting them.
- Policy-based judges: your own content policy (for example "no investment advice", "no medical dosage") as binary rubrics.
- Adversarial datasets: red-team prompts, jailbreak attempts and sensitive-topic questions (the AI Security course covers attack methods in depth).
Refusals: both directions
Measure two rates on labeled sets:
- Harmful-request refusal rate: should refuse or safely redirect.
- Over-refusal rate: benign requests wrongly refused ("How do I kill a Python process?").
Tuning safety prompts often trades one for the other; track both.
Privacy and PII leakage
- Code-based detectors (regex for emails, phone numbers, national ID formats such as Pakistan's CNIC or UAE Emirates ID patterns, card numbers with Luhn checks) plus NER-based PII tools such as Microsoft Presidio.
- Adversarial prompts that try to extract other users' data or system prompts.
- Checks on logs and traces too: PII in telemetry is a leak.
Fairness and consistency across groups
Run the same questions varying only a demographic or regional attribute (name, country, language) and compare outcomes: counterfactual testing. For example, does a loan-information assistant give systematically different information when the user writes in Urdu versus English? Differences are not always bias, but they must be explainable.
Tone and brand
Binary criteria derived from your brand guide ("uses customer's name once", "no exclamation marks in complaints", "British spelling") or pairwise comparison against approved reference answers.
Worked example: a health information chatbot in the UAE
A clinic group's assistant answered appointment and general wellness questions. Their safety suite: 300 adversarial prompts (dosage requests, emergency symptoms, requests for diagnoses), with a code check that emergency keywords trigger the emergency-contact response, a policy judge for "no diagnosis or dosage", over-refusal tests on benign questions, groundedness on clinic information, and PII detection on outputs and logs. Emergency handling was a blocking criterion at 100%.
Pitfalls
- Treating moderation scores as complete safety. They miss domain-specific harms.
- Ignoring over-refusal, which frustrates users and pushes them elsewhere.
- English-only safety tests for multilingual products.
- Forgetting telemetry as a privacy surface.
How to measure success
A dashboard with groundedness, harmful-refusal, over-refusal, PII-leak and counterfactual-consistency rates per release, with blocking thresholds for must-never-happen failures.
Key takeaways
- Groundedness: decompose into claims, verify each against context, report unsupported claims.
- Layer safety: moderation classifiers (check language coverage), policy judges and adversarial sets.
- Track harmful-request refusal and over-refusal together.
- Detect PII in outputs and telemetry; run counterfactual tests for fairness; encode brand tone as criteria.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Add an over-refusal set of 30 benign prompts and a PII detector for outputs and traces to your eval suite, and report both rates.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.