---
title: "Safety, limits and verification | Optimize All Academy"
description: "What can go wrong Claude is among the most carefully built AI systems available, but every language model shares some limits: - Hallucination: plausible…"
url: https://optimizeall.com/learn/mastering-claude/safety-limits-and-verification
updated: 2026-10-05
---

Mastering Claude (Anthropic) · Safety, privacy and team adoption · lesson 18 of 20 · 14 min

# Safety, limits and verification

## What can go wrong

Claude is among the most carefully built AI systems available, but every language model shares some limits:

- **Hallucination:** plausible but false or unsupported statements, including invented citations, quotes and figures.
- **Knowledge cut-off:** training data stops at a date; without search or sources, recent facts are missing.
- **Reasoning and arithmetic slips:** much rarer with thinking and code execution, but not zero, especially with wrong inputs.
- **Over-confidence and sycophancy:** agreeing with a flawed premise, or softening bad news. Ask explicitly for disagreement.
- **Misreading inputs:** scans, charts, ambiguous column headers.
- **Manipulated inputs:** prompt injection in web pages, emails and documents.

Anthropic publishes system cards and a usage policy and trains Claude with safety methods (including its constitution-based approach), and newer models run safeguards that may decline certain high-risk requests. None of this removes your responsibility for what you publish or decide.

## The verification ladder, scaled to stakes

| Stakes | Example | Minimum verification |
|---|---|---|
| Low | Brainstorming post ideas | Read it; use judgement |
| Medium | Client email, blog post | Check facts, names, numbers, links |
| High | Contract summary, pricing, regulatory claims | Source-check every material claim; second person reviews |
| Critical | Legal, medical, tax, financial or safety decisions | Qualified human expert decides; AI is input only |

Techniques that raise reliability:

- Ask for **quotes and references** from supplied documents.
- Ask **"what would make this wrong?"** and **"which claims are you least sure about?"**
- Use **code execution** for calculations, and check two numbers yourself.
- Run important questions **twice** or in two tools; disagreement signals where to look.
- Keep a **"checks needed"** section in every output format.

## Responsible marketing and content

- **Never fabricate** reviews, testimonials, case-study results, awards or quotes. In the US, the FTC's rule on fake reviews and testimonials allows civil penalties; in the UK, the Digital Markets, Competition and Consumers Act 2024 makes fake reviews a banned practice, and the ASA/CAP Code requires substantiation of claims.
- **Disclose** paid partnerships and material connections per your markets' rules (for example #ad or "Paid partnership" in the UK, platform tools, and local regulator guidance in the UAE and KSA).
- **Label AI-generated realistic imagery or media** where platforms or laws require.
- **Respect intellectual property**; do not ask Claude to reproduce copyrighted works wholesale.

## When Claude declines

If Claude declines a legitimate request, give honest context about your purpose ("I run security-awareness training for our staff; I need examples of phishing red flags"). Rephrasing to describe the real goal often helps. Do not try to jailbreak or trick the model: it breaches usage policies, can get accounts suspended, and produces less reliable outputs. For API builders, check `stop_reason` for refusals and route those cases to a human or a documented fallback.

## Worked example: the invented statistic

A junior marketer asks Claude for "stats on why brands should use micro-influencers". The draft includes a precise percentage attributed to "a 2025 industry study". She asks: "Which source is this from? Give the URL." Claude cannot provide one and says the figure should be treated as unverified. She replaces it with a qualitative claim and a real, cited survey she finds herself. The lesson: *every specific number needs a source you have opened*.

## A reusable self-check prompt

Add this to the end of any high-stakes request, or save it as a follow-up you paste after the draft arrives:

```text
Before I use this, audit your own answer:
1. List every specific number, date, name, quote and citation.
2. For each, say where it came from: the material I gave you, a source you
   searched (with link), or your general knowledge.
3. Mark anything from general knowledge or without a source as UNVERIFIED.
4. Name the three claims most likely to be wrong and why.
```

The audit does not replace opening sources, but it tells you exactly which lines to check first.

## Hands-on: build your verification habit

1. Take one AI-assisted piece of work from this month.
2. Classify its stakes using the table.
3. Apply the matching verification and record what you checked.
4. Add a "Checks needed" section to your most-used prompt or Project instructions.

## Pitfalls

- Asking Claude whether its own citation is real and accepting "yes" without opening it.
- Treating fluent, formatted output as evidence of accuracy.
- Letting urgency skip the ladder for high-stakes content.

## How to measure success

No unverified number or citation reaches publication, high-stakes outputs always get a second reviewer, and your error log shows errors caught before release rather than after.

## Video lecture: Safety, limits and verification

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Safety, limits and verification
2. Why this is your job
3. Shared limits
4. The verification ladder
5. Reliability techniques
6. Marketing red lines
7. When Claude declines
8. Simple example
9. Worked example: the invented statistic
10. Pitfalls
11. Try this now
12. Watch me do it, part 1
13. Watch me do it, part 2
14. Recap and next step

## Lecture transcript

### Safety, limits and verification

The most dangerous AI output is not the obviously wrong one. It is the polished, confident, perfectly formatted paragraph with one invented number in it. In this lecture you will learn the limits every language model shares, a verification ladder that scales with the stakes, the marketing rules you must never break, and what to do when Claude declines a request.

### Why this is your job

Why is verification your job and not the model's? Because fluency is not accuracy. A beautifully written paragraph can contain one invented figure, and it will look exactly as confident as the true sentences around it. And when that paragraph goes to a client, it goes out under your name, not the model's. Think of AI like a very fast junior analyst. Brilliant at drafting, and you still review before anything leaves the building.

### Shared limits

Every language model shares some limits. Hallucination, meaning plausible but false statements, including invented citations and figures. A knowledge cut off. Occasional reasoning or arithmetic slips, rarer with thinking and code, but not zero. A tendency to agree with a flawed premise or soften bad news, so ask explicitly for disagreement. Misreading scans and ambiguous data. And manipulated inputs through prompt injection. Anthropic invests heavily in safety, and newer models run safeguards that may decline some high risk requests. But none of that transfers your responsibility to the model.

### The verification ladder

Scale your checking to the stakes. Low stakes, like brainstorming post ideas, just read it and use judgement. Medium, like a client email or blog post, check facts, names, numbers and links. High, like a contract summary, pricing or a regulatory claim, source check every material claim and get a second person to review. Critical, like legal, medical, tax, financial or safety decisions, a qualified human expert decides, and the AI is only an input.

### Reliability techniques

Some techniques raise reliability dramatically. Ask for quotes and references from supplied documents. Ask, what would make this wrong, and which claims are you least sure about. Use code execution for calculations, and still check two numbers yourself. Run important questions twice or in two different tools, because disagreement shows you exactly where to look. And add a checks needed section to every output format, so uncertainty is always visible.

### Marketing red lines

Now the red lines for marketers. Never fabricate reviews, testimonials, case study results, awards or quotes. In the United States, the FTC's rule on fake reviews allows civil penalties, and in the United Kingdom, fake reviews are a banned practice under the Digital Markets, Competition and Consumers Act. Claims must be substantiated under the advertising codes. Disclose paid partnerships as your markets require. Label realistic AI generated media where platforms or laws require it. And respect intellectual property.

### When Claude declines

Sometimes Claude declines a request you consider legitimate. Give honest context. I run security awareness training for our staff and need examples of phishing red flags. Describing the real goal usually unlocks a helpful answer. Do not try to jailbreak or trick the model. It breaches the usage policy, can get accounts suspended, and produces less reliable output anyway. If you build on the API, check the stop reason for refusals and route those cases to a person or a documented fallback.

### Simple example

A simple example. Ask Claude without search for three sources supporting a claim about email open rates. You receive three neatly formatted citations. Open each one. Two exist and say roughly what was claimed. The third cannot be found anywhere. It was generated to look right. Delete it, and if you need three sources, run the question with web search on and open those too. The lesson is not that Claude is unreliable. It is that citations without search, or without opening, are not evidence.

### Worked example: the invented statistic

A junior marketer asks for statistics on why brands should use micro influencers. The draft includes a precise percentage attributed to a twenty twenty five industry study. She asks which source, and for the link. Claude cannot provide one and says to treat the figure as unverified. She deletes it, writes a qualitative claim instead, and adds a real survey she found and read herself. The rule is simple. Every specific number needs a source you have personally opened. Afterwards, she added a rule to her team's Project instructions. Any statistic must include its source name and a link, or be written as a qualitative statement. Over the next month, the checks needed lists at the end of each draft became shorter, because fewer unsourced numbers were appearing in the first place. The best verification is the kind you build into the prompt.

### Pitfalls

Three pitfalls. Asking Claude whether its own citation is real and accepting yes without opening it. Treating fluent, formatted output as evidence of accuracy. And letting urgency skip the ladder for high stakes content, which is exactly when mistakes are most expensive.

### Try this now

Try this now. Find one piece of AI assisted work you produced this month, a client email, a report, a post. Decide its stakes using the ladder, low, medium, high or critical. Then apply the matching checks. For medium, check every name, number and link. For high, open the source of every material claim and ask a colleague to review. Paste the self audit prompt from the lesson and see what it flags. Finally, add a checks needed line to the prompt or Project you use most, so every future output tells you what to verify.

### Watch me do it, part 1

Let me verify a real kind of deliverable, a client report on micro influencer performance. First I classify the stakes. It goes to a client and includes numbers they will use for budget decisions, so it is high. Next I paste the self audit prompt. List every number, date, name, quote and citation, say where each came from, the material I gave you, a searched source with a link, or general knowledge, mark anything without a source as unverified, and name the three claims most likely to be wrong. The audit lists fourteen specifics. Eleven trace to the client's own export. Three are marked unverified, including an industry engagement benchmark.

### Watch me do it, part 2

Now I deal with the three unverified items. The engagement benchmark has no primary source I can find, so I delete it and write a qualitative statement instead. The date of a platform policy change I find on the platform's official page, so it turns green. The third is a quote attributed to the client's head of marketing, taken from a call summary, so I send it to her for approval rather than guessing. Because the report is high stakes, a colleague checks three random figures against the export and initials it. Twenty minutes of checking, and nothing in the report is invented.

### Recap and next step

Recap. Know the shared limits. Scale verification to stakes. Hold the marketing red lines. Give honest context when Claude declines, never jailbreak. Your next step: take one AI assisted piece from this month, classify its stakes, apply the matching checks, and add a checks needed section to your most used prompt or Project instructions.

## Key takeaways

- All language models can hallucinate, miss recent facts, slip on arithmetic, agree too readily and be manipulated by injected content.
- Scale verification to the stakes, up to qualified human decisions for legal, medical, tax, financial or safety matters.
- Never fabricate reviews or testimonials; substantiate claims, disclose partnerships and label realistic AI media where required.
- If Claude declines a legitimate task, give honest context; never jailbreak. API builders should handle refusals in code.

## Try it

Take one AI-assisted piece of work from this month and apply the verification ladder. Note its stakes level and what you checked, and record any error you found.

- [Previous: Integrations: connect Claude to your CRM, forms, Slack and automation tools](https://optimizeall.com/learn/mastering-claude/claude-api-business-integrations)
- [Next: Privacy and data controls](https://optimizeall.com/learn/mastering-claude/privacy-and-data-controls)
- [All lessons of Mastering Claude (Anthropic)](https://optimizeall.com/learn/mastering-claude)
