Skip to content

Prompt Engineering for Content & Sales · Brand voice, evaluation and originality · lesson 13 of 15 · 12 min

Evaluating outputs and fact-checking before you publish

Quality is a process, not a feeling

"It looks good" is not an evaluation. Professionals use a consistent checklist so quality does not depend on mood or deadline pressure. This lesson gives you a practical evaluation rubric and a fact-checking routine for AI-assisted content.

The CLEAR rubric

Score each piece from 1 to 5 on:

  • C: Correct. Are facts, numbers, names, prices, dates and claims verified?
  • L: Legal and compliant. Disclosures (#ad, paid partnership, AI labels where required), platform ad policies, no misleading claims, no infringing content.
  • E: Engaging. Does the hook earn attention? Is there a specific detail or story?
  • A: Audience fit. Right language, culture, sophistication and platform norms?
  • R: Recognisably us. Does it match the voice guide and sound human, not generic?

Anything scoring 1 or 2 on C or L does not ship, regardless of how engaging it is.

Fact-checking routine

  1. Extract claims. Ask the model: "List every factual claim in this text as a bullet list." That makes checking faster.
  2. Categorize. Brand facts (price, ingredients, dates) are checked against the client brief or product page. External facts (market data, news, rules) are checked against primary sources: the regulator, the official report or the platform's help center.
  3. Verify or cut. If you cannot verify within reasonable time, rewrite qualitatively or remove.
  4. Watch for subtle drift. AI may turn "up to 40% off selected items" into "40% off everything". Compare offer wording word for word.
  5. Record sources for sponsored and regulated content.

Using AI to help evaluate

AI can support, but not replace, your evaluation:

  • "Evaluate this caption against the CLEAR rubric and our voice guide. Be critical and specific."
  • "Identify any statements that could be seen as misleading to a consumer."
  • "Which parts of this would a native Gulf Arabic speaker find unnatural?" This is a pointer only; confirm with a native speaker.

Be aware that models can be agreeable and may rate their own work generously. Use a fresh chat for evaluation, and ask for weaknesses specifically.

A/B testing: letting the audience decide

For hooks, subject lines and ads, the best evaluator is your audience. Use AI to generate distinct variants, not near-duplicates, then test them with platform A/B tools or by posting variants over time. Keep a simple log: variant, angle, result. Feed winners back into your examples.

Hands-on: a CLEAR review routine in a fresh chat

Run evaluation in a fresh chat (or a separate "Reviewer" project) so the model isn't grading its own conversation:

You are a strict reviewer. Evaluate the draft below against:
- Our voice guide (attached)
- The CLEAR rubric: Correct, Legal/compliant, Engaging, Audience fit, Recognizably us (score 1-5 each)
For C and L, list every claim, number, offer term and disclosure, and say whether it matches the attached
offer sheet / approved claims (match / mismatch / not in sources).
Be specific and critical. Do not rewrite; list issues in order of severity.
[paste draft]

Before: "Looks good!" from the same chat that wrote the caption.

After: a scored review that flags "40% off everything" as a mismatch with the offer sheet ("up to 40% off selected lines"), a missing #ad, and one banned word, in order of severity. You fix them, then a human approves.

Build a simple test log

| Date | Asset | Variant | Angle | Result (metric) | Keep/kill | |---|---|---|---|---|---|

Feed winning variants back into your example bank. Over time, your templates learn from your audience, not just from you.

Worked example: a sponsored post review

A US-based fitness creator receives an AI-assisted caption for a protein bar sponsorship:

"Game-changing protein bar with 20g protein that melts fat while you sleep! Use my code for 50% off"

CLEAR review:

  • Correct: protein amount confirmed from the pack. "Melts fat while you sleep" is unsupported. Discount is actually 20%. Score 1.
  • Legal: no sponsorship disclosure; a health claim the brand did not approve; wrong discount. Score 1.
  • Engaging: decent energy. 3.
  • Audience fit: OK. 3.
  • Recognisably us: "game-changing" is on the ban list. 2.

Revised: "#ad I've been keeping these in my gym bag for post-workout snacks: 20g protein and they actually taste like dessert. Code SAM20 gets you 20% off." Plus the platform's paid-partnership label.

Pitfalls

  • Skipping review "just this once" on a busy day.
  • Checking only the first draft, not the final revised version.
  • Treating AI evaluation as final approval.

Video lecture: Evaluating outputs and fact-checking before you publish

Lecture coming soon · 12 chapters · about 8 minutes. Read the full transcript below.

  1. Evaluate before you publish
  2. Why a rubric
  3. CLEAR
  4. The airport security analogy
  5. The fact-checking routine
  6. AI as reviewer
  7. Example 1: the sponsored protein bar post
  8. Example 2: the reviewer project
  9. Watch me do it
  10. Scale the review to the stakes
  11. Let the audience decide
  12. Recap and try this now

Lecture transcript

Evaluate before you publish

It looks good. Those three words have published more mistakes than any AI model. Wrong discounts, missing disclosures, health claims nobody approved, all in content that looked good. In this lecture you'll learn to replace it looks good with a real evaluation routine: the CLEAR rubric, a five step fact checking process, and a way to use AI as a strict reviewer without letting it grade its own homework. You'll see a sponsored post that nearly went wrong, and watch me run a review that catches three problems in thirty seconds.

Why a rubric

Why use a rubric? Because quality shouldn't depend on your mood, or on how close the deadline is. Deadlines are exactly when errors slip through, because checking feels like a luxury. A rubric makes the check automatic and consistent: the same five questions, every time, for everyone. And it makes one rule clear. Some problems block publishing no matter how engaging the content is. A brilliant caption with a false claim or a missing disclosure doesn't ship.

CLEAR

Here's the rubric. Score each piece from one to five. C, correct: are facts, numbers, names, prices, dates and claims verified? L, legal and compliant: disclosures like hashtag ad and paid partnership labels, AI labels where required, platform ad policies, and no misleading claims. E, engaging: does the hook earn attention, is there a specific detail or story? A, audience fit: the right language, culture and platform norms? And R, recognizably us: does it match the voice guide and sound human? Anything scoring one or two on correct or legal doesn't ship, however engaging it is.

The airport security analogy

Think of airport security. Everyone goes through the scanner, even frequent flyers, even people running late for their flight. Some items are never allowed through, no matter how nice the passenger is. Your CLEAR review is the scanner, and false claims, missing disclosures and misleading offers are the prohibited items. The most dangerous moment is when you're in a hurry and think, just this once. That's exactly when the scanner matters most.

The fact-checking routine

Now the fact checking routine. One, extract the claims: ask the model to list every factual claim as a bullet list. Two, categorize. Brand facts, like price, ingredients and dates, are checked against the brief, offer sheet or product page. External facts, like market data or rules, are checked against primary sources, like the regulator or the official report. Three, verify or cut: if you can't verify it in reasonable time, rewrite it qualitatively or remove it. Four, watch for subtle drift. AI can turn up to forty percent off selected items into forty percent off everything. Compare offer wording word for word. And five, record sources for sponsored and regulated content.

AI as reviewer

Can AI help evaluate? Yes, as a pointer, not a verdict. Models can be agreeable, and they tend to rate their own work generously, especially in the same conversation that produced it. So run evaluation in a fresh chat, or a separate reviewer project. Tell it to be a strict reviewer, score against CLEAR and your voice guide, check every claim and offer term against the attached offer sheet and approved claims, and list issues in order of severity without rewriting. Then a human makes the call. For language nuance, like whether Gulf Arabic sounds natural, the AI can flag concerns, but a native speaker confirms.

Example 1: the sponsored protein bar post

Here's a real style example. A US fitness creator receives an AI assisted caption for a protein bar sponsorship: game changing protein bar with twenty grams of protein that melts fat while you sleep, use my code for fifty percent off. The CLEAR review. Correct: the protein amount checks out, but melts fat while you sleep is unsupported, and the discount is actually twenty percent. Score one. Legal: no sponsorship disclosure, an unapproved health claim, a wrong discount. Score one. Engaging: three. Audience: three. Recognizably us: game changer is on the ban list. Two. The revision: hashtag ad, keeping these in my gym bag for post workout snacks, twenty grams of protein and they taste like dessert, code for twenty percent off, plus the paid partnership label.

Example 2: the reviewer project

Now a team process, with illustrative details. A marketing team at an electronics retailer in Riyadh sets up a separate reviewer project containing only the CLEAR prompt, the voice guide, the current offer sheet and approved claims. Every asset goes through it before a human approves. In the first month, the reviewer catches a string of offer mismatches: wrong end dates, a missing selected items qualifier, and a warranty length that changed last quarter. Nothing publishes without a human sign off. But the human now spends their time on judgment calls, not on hunting for typos in discount terms.

Watch me do it

Let me run a review. I open a fresh chat, not the one that wrote the draft. I attach the voice guide and the offer sheet, and paste the reviewer prompt from the lesson text, then an email promoting a furniture sale. Thirty seconds later: a CLEAR scorecard and a claims table. Offer: forty percent off everything. Offer sheet: up to forty percent off selected lines. Mismatch. End date: Sunday. Offer sheet: Sunday. Match. Free delivery: not in sources. And in the severity list: the offer mismatch first, the unverified free delivery second, and a banned word third. I fix all three, and a colleague approves it. The same chat that wrote the email had called it ready to send.

Scale the review to the stakes

You don't need the same depth of review for everything. Scale it to the stakes. A behind the scenes post with no claims and no offer: a quick CLEAR skim. A promotional email with prices and dates: the full review, with every offer term checked against the offer sheet. And anything with health, financial or legal claims, sponsored content in regulated categories, or content aimed at children: a full review plus sign off from the right expert or your legal team. The rubric stays the same. What changes is how hard you look, and who else needs to look.

Let the audience decide

One more evaluator: your audience. For hooks, subject lines and ads, generate genuinely distinct variants, not near duplicates, and test them with platform A B tools or by posting variants over time. Keep a simple test log: date, asset, variant, angle, result, keep or kill. And feed the winners back into your example bank. Over time, your templates learn from your audience, not just from your taste.

Recap and try this now

Let's recap. Evaluate everything with CLEAR: correct, legal, engaging, audience fit, recognizably us, and never ship low scores on correct or legal. Fact check by extracting claims, checking brand facts against your sources and external facts against primary sources, and watching for drift. Use AI as a strict reviewer in a fresh chat or reviewer project, but keep human approval. And let the audience decide between distinct variants. Here's your try this now. Set up a reviewer project with the prompt from the lesson text, your voice guide and offer sheet, and run your next three assets through it before publishing.

Key takeaways

  • Use a consistent rubric (CLEAR: Correct, Legal, Engaging, Audience fit, Recognizably us); low C or L scores block publishing.
  • Fact-check by extracting claims, categorizing brand vs external facts, and verifying against the source; watch for subtle drift.
  • Evaluate in a fresh chat or reviewer project, ask for weaknesses, and treat AI evaluation as a pointer, not approval.
  • Let the audience decide between distinct variants and log results back into your example bank.

Try it

Score your last three published AI-assisted posts with the CLEAR rubric and note one process change that would have raised the lowest score.