Evaluating and Monitoring LLM ApplicationsMetrics and graders: code, judges and quality dimensions · Lesson 6 of 16

Groundedness, safety, privacy and fairness metrics

Article · 14 min · 8 min lecture

Video lecture

Groundedness, safety, privacy and fairness metrics

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Groundedness, safety and quality metrics

  • Groundedness
  • Safety and refusals
  • Privacy
  • Fairness and tone

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Beyond "is it correct?"

Production LLM systems need evaluation on several dimensions that are easy to forget until something goes wrong publicly: groundedness (is it supported by the provided sources?), safety (toxicity, harmful advice, policy violations), privacy (PII leakage), refusal behavior (refusing when it should, not refusing when it should not), bias and fairness across user groups, and brand and tone. Each needs specific datasets and graders.

Groundedness and faithfulness

A response is grounded (or faithful) when every factual claim is supported by the context provided (retrieved documents, tool outputs). Ungrounded claims, often called hallucinations, are the top failure of RAG and support systems.

A common method, used by open-source libraries such as Ragas and DeepEval in various forms:

  1. Decompose the answer into atomic claims.
  2. Verify each claim against the context (supported / not supported).
  3. Score = supported claims / total claims.
# groundedness.py: claim-level groundedness with an LLM (sketch; add retries and logging)
import json, os
from anthropic import Anthropic
client = Anthropic(); MODEL = os.environ["JUDGE_MODEL"]

def ask_json(prompt: str) -> dict:
    msg = client.messages.create(model=MODEL, max_tokens=800, temperature=0,
                                 messages=[{"role": "user", "content": prompt}])
    t = msg.content[0].text
    return json.loads(t[t.find("{"): t.rfind("}") + 1])

def groundedness(answer: str, context: str) -> dict:
    claims = ask_json("Split the ANSWER into atomic factual claims. Ignore greetings and "
                      "questions. Return {\"claims\": [..]}.\nANSWER:\n" + answer)["claims"]
    results = []
    for c in claims:
        r = ask_json("Is the CLAIM fully supported by the CONTEXT? Answer only from the context. "
                     "Return {\"supported\": true|false, \"evidence\": \"<quote or empty>\"}.\n"
                     f"CONTEXT:\n{context}\n\nCLAIM:\n{c}")
        results.append({"claim": c, **r})
    score = sum(r["supported"] for r in results) / len(results) if results else 1.0
    return {"score": score, "unsupported": [r["claim"] for r in results if not r["supported"]]}

Calibrate this like any judge (previous lesson). Watch for claims that come from general knowledge: decide whether your product allows them.

Safety and toxicity

  • Classifier-based: use moderation endpoints or open safety classifiers (for example the moderation APIs offered by major providers, or open models such as Llama Guard-style classifiers) to flag categories like hate, harassment, self-harm, sexual content and violence. Check category definitions and language coverage for Arabic and Urdu before trusting them.
  • Policy-based judges: your own content policy (for example "no investment advice", "no medical dosage") as binary rubrics.
  • Adversarial datasets: red-team prompts, jailbreak attempts and sensitive-topic questions (the AI Security course covers attack methods in depth).

Refusals: both directions

Measure two rates on labeled sets:

  • Harmful-request refusal rate: should refuse or safely redirect.
  • Over-refusal rate: benign requests wrongly refused ("How do I kill a Python process?").

Tuning safety prompts often trades one for the other; track both.

Privacy and PII leakage

  • Code-based detectors (regex for emails, phone numbers, national ID formats such as Pakistan's CNIC or UAE Emirates ID patterns, card numbers with Luhn checks) plus NER-based PII tools such as Microsoft Presidio.
  • Adversarial prompts that try to extract other users' data or system prompts.
  • Checks on logs and traces too: PII in telemetry is a leak.

Fairness and consistency across groups

Run the same questions varying only a demographic or regional attribute (name, country, language) and compare outcomes: counterfactual testing. For example, does a loan-information assistant give systematically different information when the user writes in Urdu versus English? Differences are not always bias, but they must be explainable.

Tone and brand

Binary criteria derived from your brand guide ("uses customer's name once", "no exclamation marks in complaints", "British spelling") or pairwise comparison against approved reference answers.

Worked example: a health information chatbot in the UAE

A clinic group's assistant answered appointment and general wellness questions. Their safety suite: 300 adversarial prompts (dosage requests, emergency symptoms, requests for diagnoses), with a code check that emergency keywords trigger the emergency-contact response, a policy judge for "no diagnosis or dosage", over-refusal tests on benign questions, groundedness on clinic information, and PII detection on outputs and logs. Emergency handling was a blocking criterion at 100%.

Pitfalls

  • Treating moderation scores as complete safety. They miss domain-specific harms.
  • Ignoring over-refusal, which frustrates users and pushes them elsewhere.
  • English-only safety tests for multilingual products.
  • Forgetting telemetry as a privacy surface.

How to measure success

A dashboard with groundedness, harmful-refusal, over-refusal, PII-leak and counterfactual-consistency rates per release, with blocking thresholds for must-never-happen failures.

Key takeaways

  • Groundedness: decompose into claims, verify each against context, report unsupported claims.
  • Layer safety: moderation classifiers (check language coverage), policy judges and adversarial sets.
  • Track harmful-request refusal and over-refusal together.
  • Detect PII in outputs and telemetry; run counterfactual tests for fairness; encode brand tone as criteria.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Tightening a safety prompt reduces harmful answers but users complain about refusals to ask how to 'kill a stuck process'. What metric was missing?
  2. How is claim-level groundedness typically computed?

Put it into practice

Add an over-refusal set of 30 benign prompts and a PII detector for outputs and traces to your eval suite, and report both rates.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.