Off-Page SEO & Link BuildingDigital PR and AI-assisted link workflows · Lesson 13 of 17
AI-assisted link-building workflows with human QA
Video lecture
AI-assisted link-building workflows with human QA
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 AI-assisted link building
AI can read a hundred web pages while you finish your coffee. So why aren't link builders simply automating the whole job? Because the parts AI is fast at are not the parts that decide whether a link is safe or a pitch is welcome. In this lecture you'll learn where AI genuinely helps in link building, where it must never decide, and how to build a small triage pipeline with human quality control baked in.
0:33 Why it matters
Why does this matter now? Because AI has made it nearly free to generate thousands of personalised sounding emails. Editors are drowning in them. Teams that use AI to send more are burning their sender reputation and their brand. Teams that use AI to research better, and then send fewer, sharper pitches, are winning. The tool is the same. The strategy decides the outcome.
1:01 The key idea
Here's the key idea, with an analogy. Think of AI as a brilliant research intern on their first day. They read fast, summarise well and draft decently. But they don't know your clients, they can't tell which blogs secretly sell links, and occasionally they'll confidently make something up. You'd never let a first day intern sign contracts or email your best journalist contacts unsupervised. Same rules here. AI prepares. Humans decide and send.
1:33 Where AI helps
So where does AI help? Five places. Prospect triage: summarising a site's topics, classifying the page type, flagging write for us, all topics accepted menus. Personalisation research: summarising a journalist's last five articles into beats and angles. Drafting: turning your segment template and notes into a body paragraph. Coverage tracking: extracting outlet, date, whether you're linked, and which finding they quoted. And asset support: cleaning data, drafting captions and translating summaries.
2:04 AI must not decide
And where must it not decide? Whether a site is safe. Link sellers hide their menus, and only human review and evidence decide. Facts, findings and quotes. Models can invent plausible numbers and compliments. Volume. More emails is not a strategy. And personal data. Don't send contacts' personal details to an AI provider unless your agreements and privacy notice cover it. Prefer sending only public page text.
2:33 The triage pipeline
Let's walk through the triage pipeline in the lesson. It reads a CSV of prospect URLs. For each one, it fetches the page, strips scripts, navigation and footers, and keeps the visible text. It sends that text to an LLM with a system prompt that says: use only the page text, return JSON with page type, main topics, whether it accepts paid posts, and evidence quotes, and if something isn't shown, say unknown. The result goes into an AI notes column. And there's a blank human decision column, waiting for a person.
3:13 Engineering habits
A few engineering habits matter. Keys live in environment variables, never in code. The model ID comes from an environment variable too, set from the provider's current models list, so you're not hard coding something that changes. Errors are caught per row, so one broken page doesn't stop the batch. And you log the model and date with the output, so you can reproduce results later. Start with twenty pages. Compare the AI notes with your own judgement before you scale.
3:48 Example 1 (simple)
Worked example one, simple. A solo consultant has forty resource pages to check for a client's free guide. She runs them through the prompt manually in her AI assistant, one page's text at a time. The notes quickly show that eleven are really guest post menus and six haven't been updated in years. She reads the remaining twenty-three herself and pitches nine. The AI saved her an afternoon of skimming, but every final decision was hers.
4:21 Example 2 (illustrative)
Worked example two, realistic and illustrative. A small agency in Islamabad qualifies about a hundred and fifty new prospects a week for three clients. With the triage script, reviewers read AI notes first and open full pages only for plausible candidates. Sampling shows the model reliably spots guest post menus, but sometimes labels news round ups as resource lists. So they add an example of each to the system prompt. Agreement improves. And instead of sending more emails, they reinvest the saved hours into sharper first lines and relationship follow ups.
5:01 Test your AI notes
Let's look at how to check whether your AI notes are trustworthy. Take twenty pages. Classify them yourself first, without looking at the AI output. Then compare. Count how often you agree on page type and on whether the site accepts paid posts. If you disagree on more than, say, two in twenty, look at the disagreements. Often the model lacked context, like the difference between a news round up and a curated resource list. Add one short example of each to the prompt and re-test. That's prompt improvement based on evidence, not vibes.
5:42 Watch me do it: pilot the triage script
Watch me do it. I'll pilot the triage script on twenty prospects before trusting it with more. First, I set two environment variables in my terminal: the API key and the model ID I've picked from the provider's current model list. Nothing goes into the code. Next, I open prospects dot CSV and make sure it has a url column with twenty rows. Then, before running anything, I classify the twenty pages myself in a separate column, page type and paid posts yes, no or unknown. That's my ground truth. Now I run the script. It fetches each page, sends only the visible text, and writes the model's JSON into an ai notes column. Two rows show errors because the pages blocked the request, and that's fine, they're logged. I compare the results with mine. We agree on sixteen of eighteen. Both disagreements are news round ups the model called resource lists. So I add one short example of each to the system prompt, and rerun on those rows. Now they match. Finally, I note the model ID, date and agreement rate at the top of the output file. Only now would I let the team use it on a hundred prospects, with humans still deciding.
7:12 Drafting and controls
Now drafting. The lesson includes a prompt that asks the model to write the body only, maximum ninety words, using only the facts you provide, and to write a MISSING marker when a fact is absent, instead of inventing one. That one rule prevents fabricated statistics and fake compliments. Then a human writes the first line from real research and reads the whole email before sending. Add three controls: weekly sampling of AI classifications, blocklists that always win, and no auto send, ever.
7:48 Common mistakes + recap
Common mistakes. Letting AI decide which sites are safe. Sending AI drafts without a human first line and review. Scaling send volume just because drafting got cheaper. And pasting contacts' personal data into AI tools without the right agreements. Recap: AI summarises, classifies, extracts and drafts. Humans decide, verify and send. Structured JSON with evidence makes review fast. And prompts should forbid invention. Try this now. Run the triage prompt on twenty prospects, record your own decisions, and calculate how often you agreed.
Where AI genuinely helps
Large language models (LLMs) such as Claude, ChatGPT and Gemini are good at reading and summarising text, classifying pages into categories you define, extracting structured fields, and drafting from templates and notes. In link building that maps to:
| Task | AI role | Human role |
|---|---|---|
| Prospect triage | Summarise a site's recent topics; classify page type; flag "write for us — all topics" menus | Decide; read two articles for shortlisted sites |
| Personalisation research | Summarise a journalist's last five articles into beats and angles | Choose the angle; write the first line |
| Drafting | Draft the body of a segment template from your notes | Edit, fact-check, send |
| Coverage tracking | Extract outlet, date, whether your brand is linked, and quoted finding from an article | Spot-check accuracy |
| Asset support | Clean data, draft chart captions, translate summaries | Verify every number and quote |
Where AI must not decide
- Whether a site is safe to work with. Link sellers hide their menus; only human review and evidence decide.
- Facts, findings and quotes. Models can invent plausible numbers and compliments. Every claim in an email must come from your data or your notes.
- Volume. AI makes it cheap to send thousands of "personalised" emails. That is spam with better grammar, and it will damage your sender reputation and your brand.
- Personal data. Do not send contacts' personal data to an AI provider unless your organisation's agreements and privacy notice cover it; prefer sending only public page text.
Hands-on: a prospect-triage pipeline in Python
This script fetches each prospect page, extracts visible text, and asks an LLM to return structured JSON that goes into your sheet as notes for a human. It uses the Anthropic Python SDK; the same pattern works with other providers' SDKs. Set ANTHROPIC_API_KEY and ANTHROPIC_MODEL (a current model ID from the provider's docs) as environment variables — never hard-code keys.
import csv
import json
import os
import sys
import anthropic
import requests
from bs4 import BeautifulSoup
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
MODEL = os.environ["ANTHROPIC_MODEL"] # e.g. set from the current models list
SYSTEM = (
"You classify web pages for a link-building team. Use ONLY the page text provided. "
"Return JSON with keys: page_type (one of resource_list, news_article, blog_post, "
"guest_post_menu, product_page, other), main_topics (list of up to 5 strings), "
"accepts_paid_posts (true, false or \"unknown\"), evidence (short quotes from the text). "
"If the text does not show something, use \"unknown\". Do not guess."
)
def page_text(url: str) -> str:
r = requests.get(url, timeout=15, headers={"User-Agent": "prospect-triage/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
for tag in soup(["script", "style", "nav", "footer"]):
tag.decompose()
return " ".join(soup.get_text(" ").split())[:12000]
def classify(url: str) -> dict:
text = page_text(url)
msg = client.messages.create(
model=MODEL,
max_tokens=600,
system=SYSTEM,
messages=[{"role": "user", "content": f"URL: {url}\n\nPAGE TEXT:\n{text}"}],
)
raw = msg.content[0].text.strip()
start, end = raw.find("{"), raw.rfind("}")
return json.loads(raw[start:end + 1])
rows = []
with open("prospects.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
try:
result = classify(row["url"])
row["ai_notes"] = json.dumps(result, ensure_ascii=False)
except (requests.RequestException, anthropic.APIError, ValueError) as e:
row["ai_notes"] = f"error: {e}"
print(f"{row['url']}: {e}", file=sys.stderr)
row["human_decision"] = "" # filled in by a person
rows.append(row)
with open("prospects_triaged.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=list(rows[0].keys()))
w.writeheader()
w.writerows(rows)Run it on a small batch first, compare the model's notes with your own judgement on 20 pages, and only then scale. Log the model ID and date with the output so results are reproducible.
A drafting prompt that forbids invention
You draft outreach emails. Write the BODY ONLY (no greeting, no first line, no signature),
max 90 words, British English, plain text.
Use ONLY these facts: {{facts}}
Segment template: {{template}}
Rules: do not invent statistics, compliments, article titles or claims about the recipient.
If a needed fact is missing, write [MISSING: what] instead.The [MISSING: …] convention makes gaps visible instead of letting the model fill them with fiction. A human writes the first line from real research and reads the whole email before it is sent.
Quality controls that keep AI honest
- Sampling: a reviewer checks 10% of AI-classified rows each week; if disagreement exceeds a threshold you set (for example 1 in 10), tighten the prompt or stop using AI for that task.
- Blocklists win: a domain on your blocked list is never re-qualified by AI.
- No auto-send: AI drafts go to a queue; nothing is sent without a human click.
- Disclosure where expected: never present AI-generated text as a named expert's quote.
Worked example (illustrative)
A small agency in Islamabad qualifies about 150 new prospects a week for three clients. With the triage script, reviewers read AI notes first and open full pages only for plausible candidates. They find that the model reliably spots guest-post menus and topic mixes but occasionally labels news round-ups as resource lists, so they add an example of each to the system prompt. Weekly sampling shows agreement improves; the time saved goes into better first lines and relationship follow-ups rather than more emails.
Common mistakes
- Letting AI decide which sites are safe.
- Sending AI-drafted emails without a human first line and review.
- Scaling send volume because drafting got cheaper.
- Pasting contacts' personal data into AI tools without the right agreements.
Key takeaways
- Use LLMs to summarise, classify, extract and draft — not to decide which sites are safe or to invent facts.
- Structured JSON output with an evidence field makes AI notes easy to review in a sheet.
- Forbid invention in prompts and mark missing facts explicitly; a human writes the first line and approves every send.
- Sample AI decisions weekly, keep blocklists authoritative, and protect personal data.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the triage script (or the prompt manually) on 20 prospects, record your own decision for each, and calculate how often you agreed with the AI notes. Adjust the prompt once and re-test.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.