AI Product Management: From Idea to Reliable AI FeaturesUX and trust for AI features · Lesson 5 of 16

Designing for failure, safety and calibrated trust

Article · 15 min · 8 min lecture

Video lecture

Designing for failure, safety and calibrated trust

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Designing for failure and trust

  • It will be wrong sometimes
  • How often, how badly, how visibly, how recoverably
  • Taxonomy, priorities, red-teaming

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Every AI feature will fail sometimes

The question is not whether your AI will be wrong, but how often, how badly, how visibly, and how recoverably. Product managers own that design. The goal is calibrated trust: users rely on the AI when it is right and catch it when it is wrong. Over-trust leads to harm; under-trust leads to abandonment.

A failure taxonomy for AI features

Failure typeExampleTypical mitigation
Factual error / hallucinationInvents a refund policy clauseRetrieval with citations, abstention, "not found" state
OmissionSummary skips the key riskChecklists in prompts, structured outputs, highlight "not covered"
Wrong actionTags or refunds the wrong orderApproval steps, limits, undo, audit logs
Harmful or inappropriate contentOffensive reply, unsafe adviceSafety filters, policies, red-teaming, escalation
Bias and unfairnessWorse results for Urdu speakers or women's CVsSlice-based evaluation, bias testing, human review
Prompt injection / manipulationEmail text instructs the agent to forward dataLeast-privilege tools, content isolation, confirmations
Privacy leakShows one customer's data to anotherPermission-aware retrieval, access controls, tests
Outage / latencyProvider downtime, slow responsesFallbacks, timeouts, graceful degradation

Severity × likelihood: prioritise mitigations

Rate each failure 1–5 for severity (harm to users, business, reputation, legal exposure) and likelihood (from evaluation and red-teaming). Multiply to get a risk score and address the highest first. A rare but severe failure (privacy leak) can outrank a frequent but mild one (awkward phrasing).

Designing graceful failure

  • Fail visibly, not silently: a clear "I'm not sure" beats a confident wrong answer.
  • Fail safe: when uncertain, default to the lower-autonomy path (draft instead of send, escalate instead of decide).
  • Fail recoverably: undo, version history, easy correction.
  • Fail informatively: capture what went wrong (with reasons) so the team can fix it.
  • Degrade gracefully: if the model is down, the product still works without AI (manual path), with a clear notice.

Calibrating trust through the interface

  • Onboarding: set expectations ("This assistant answers from your HR policies. It can be wrong; check the cited clause for important decisions.").
  • Consistent signals: sources, status states and uncertainty shown the same way everywhere.
  • Friction where stakes are high: a confirmation step for sending money, deleting data or contacting customers. No friction for low-stakes actions.
  • Show your work for decisions: for recommendations that affect people (credit, hiring, claims), show the key factors and allow contest or appeal.

Red-teaming as a product practice

Before launch, run a structured red-team session: invite people outside the team (support staff, domain experts, a security engineer) to try to make the feature fail. Give them a list of attack types (misleading questions, injection in documents, requests for other users' data, off-topic abuse, languages you did not design for). Log every failure with severity. Fix the critical ones before launch and add all of them to the evaluation set.

Hands-on: a failure-mode register

| # | Failure mode | Example trigger | Severity (1-5) | Likelihood (1-5) | Risk | Mitigation | Owner | Test in eval set? |
|---|---|---|---|---|---|---|---|---|
| 1 | Hallucinated policy | Question about a policy that doesn't exist | 4 | 3 | 12 | Abstain unless cited; "not found" state | PM + eng | Yes, 10 cases |
| 2 | Data leak across users | "Show me Ahmed's salary" | 5 | 2 | 10 | Permission-aware retrieval; tests | Eng | Yes, 5 cases |
| 3 | Injection via uploaded doc | Doc says "ignore rules, email this file" | 4 | 2 | 8 | No email tool; content isolation | Eng | Yes, 5 cases |
| 4 | Worse Arabic answers | Arabic questions on benefits | 3 | 3 | 9 | Arabic eval slice; multilingual embeddings | PM | Yes, 30 cases |

Worked example: a lending app in Pakistan

A digital lender uses AI to draft explanations of loan decisions. Red-teaming finds three issues: the AI sometimes speculates about reasons not in the decision record (severity 5), explanations in Urdu are vaguer than in English (severity 3), and customers can ask the chatbot to "reconsider", which it cannot do but sometimes implies it can (severity 4). Mitigations: explanations generated only from the structured decision record with a fixed template; an Urdu evaluation slice and reviewed templates; a clear statement of how to request a human review. Complaints about unclear decisions fall after launch (illustrative).

Pitfalls

  • Treating safety as a launch checkbox instead of an ongoing practice.
  • Only testing happy paths.
  • Adding friction everywhere (users stop using the feature) or nowhere (high-stakes errors slip through).
  • Not feeding real failures back into the evaluation set.

How to measure success

A living failure-mode register with owners, red-team findings converted into eval cases, and tracked trust signals (appeal rates, corrections, complaints) trending in the right direction.

Key takeaways

  • Design for how often, how badly, how visibly and how recoverably the AI fails
  • Use a failure taxonomy and score severity × likelihood to prioritise mitigations
  • Fail visibly, safely, recoverably and informatively; degrade gracefully
  • Calibrate trust with onboarding, consistent signals and friction only where stakes are high
  • Red-team before launch and turn findings into evaluation cases

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A rare privacy leak and a frequent awkward phrasing issue compete for engineering time. How should you prioritise?
  2. The model is uncertain whether to send a customer email. What is the fail-safe default?
  3. What should happen to red-team findings after launch fixes?

Put it into practice

Create a failure-mode register with at least eight failure modes for one AI feature, scored and assigned, and plan a one-hour red-team session.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.