AI Search Optimization: SEO for AI Overviews & Answer EnginesMeasuring AI visibility · Lesson 13 of 17
Automating a prompt panel with AI search APIs
Video lecture
Automating a prompt panel with AI search APIs
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Automating a prompt panel
If you've tried running a prompt panel by hand, you know the problem. Fifty prompts, three runs each, four surfaces. That's six hundred checks a month, and someone's weekend. The good news is that several providers now offer APIs with built-in web search and citations, so you can run a large panel on a schedule and log it automatically. In this lecture you'll build that pipeline, turn the log into metrics, and learn the one caveat you must put in every report.
0:36 The caveat
Here's the caveat first, because it's the most important sentence in this lesson. API answers are not the same as consumer app answers. The model, the search settings, the location, personalisation and interface features can all differ. So an API panel is a proxy. Think of it like a weather station on a hill outside town. It won't tell you exactly what's happening on your street, but it's consistent, cheap and brilliant for trends. Keep a small manual sample in the real apps to calibrate it.
1:13 Design the panel file
Design the panel before you write code. Your panel file has a stable prompt ID, the market, the language, the intent, and the prompt text. Balance your intents: recommendation, comparison, problem, brand and local. Include the languages your customers use, like Arabic for Saudi Arabia or Urdu for Pakistan. And version the file. If someone quietly rewrites twenty prompts, your trends break and you won't know why.
1:42 OpenAI Responses API + web_search
Now the OpenAI example in the lesson text. It uses the official Python SDK and the Responses API with the web search tool. You pass the model name through an environment variable, because supported models change, and you check OpenAI's docs for which ones support web search. The function sends the prompt, then walks through the response output looking for annotations of type url citation, and collects the URLs. It returns the answer text and the list of cited links.
2:17 The loop
The main loop reads the panel file, runs each prompt three times, and writes one log row per run: the date, the surface, the prompt ID, the run number, which brands were mentioned, and the cited URLs. Notice two things. The try and except block logs failures instead of crashing the whole batch at prompt forty-seven. And there's a short sleep between calls to be gentle with rate limits. The API key comes from the environment, never from the script.
2:52 Perplexity Sonar and others
Perplexity's Sonar API is OpenAI-compatible, so you can reuse the same SDK with a different base URL and your Perplexity key. Responses include citation or search result fields. Field names have changed over time, so the lesson tells you to inspect one raw response and adapt the parsing to Perplexity's current docs. Other providers, including Anthropic and Google, offer web search or grounding tools too. The pattern is always the same: call, extract the answer and cited URLs, and log.
3:27 From log to metrics
Next, metrics. A short pandas script loads the log, drops error rows, and flags whether your brand was mentioned and whether your domain was cited. Group by date and surface, and you get mention rate, citation rate and the number of runs. Add share of voice by splitting the mentioned column per brand, and a source domain pivot for your gap analysis. Now your monthly report writes itself, at least the numbers part.
3:59 Example 1: London estate agency
Worked example one, simple. A London estate agency sets up ten prompts about renting in specific boroughs, runs them weekly through one API, and emails itself a summary table. It costs little, takes an afternoon to set up, and within a month shows that answers cite a borough guide the agency wrote two years ago, with outdated rental figures. They update the guide. That's the real value: automation surfaced an accuracy problem nobody was looking for.
4:32 Example 2: Karachi fintech (illustrative)
Worked example two, with illustrative details. A Karachi fintech runs sixty prompts in English and Urdu every week through two APIs, and does a fifteen-prompt manual check in the consumer apps monthly. After a quarter, the API panel shows mention rate rising for send money to Pakistan from the UAE prompts, after a fees comparison page was rewritten answer-first. The manual sample shows the same direction. The report shows both, labels the API data as an API proxy, and lists the caveats.
5:08 Guardrails and mistakes
Guardrails. Web search tool calls are usually billed separately from tokens, so estimate the cost per run before scheduling. Keep API keys in environment variables or a secrets manager. Use official APIs rather than scraping consumer apps against their terms. Maintain brand alias lists, because substring matching misses variants and Arabic spellings. And have a human review a sample for accuracy and sentiment, because automation finds mentions, not truth. Common mistakes: presenting API results as what customers see, one run per prompt, hard-coded keys, and an unversioned prompt list.
5:47 Cadence
How often should you run it? For most businesses, weekly API runs and a monthly manual calibration is a good rhythm. Weekly gives you enough data points to see trends through the noise. Monthly manual checks keep you honest about what real users see. If you're in a fast-moving category, like travel deals or consumer finance, you might run key prompts daily around big launches or promotions. Whatever the rhythm, keep it consistent. A panel that runs whenever someone remembers is almost as bad as no panel at all, because gaps in the data look like changes in visibility.
6:30 Watch me do it: API panel from scratch
Watch me do it. I'll get the API panel running from scratch. Step one: in the terminal, I install the openai and pandas packages, and export two environment variables: my API key, and PANEL_MODEL set to a current model that supports web search, which I copied from OpenAI's docs. Step two: I create panel dot csv with ten prompts, each with an ID, market, language and intent, including two Arabic prompts. Step three: I run the script. Thirty calls later, panel log dot csv has thirty rows. One row says ERROR with a rate-limit message; the loop carried on, exactly as designed. Step four: I open the log and eyeball five rows. The mentioned column shows our brand in some runs, and the cited URLs include two domains I didn't expect. Step five: I run the metrics snippet. Mention rate and citation rate print by date and surface. Step six: I pick the same ten prompts and run them once, manually, in the consumer app, and write the results next to the API results. Seven of ten agree on whether we're mentioned. I note that in the report as the calibration, label the API numbers as an API proxy, and schedule the script to run weekly.
8:00 Reading results honestly
Let me show you how to read results without fooling yourself. Suppose your mention rate goes from thirty to forty percent in a month. Is that real? Check the number of runs behind each figure. With only thirty runs, a ten-point swing can easily be noise. Look for movement that persists across several weeks and shows up in more than one surface. Cross-check with the other evidence: did the Search Console AI impressions for the relevant page rise? Did Bing's grounding queries change? When three independent signals agree, you can say with confidence that something changed.
8:42 Recap and try this now
Recap. An API panel is a consistent, scalable proxy for trends. Design a versioned panel, call APIs with web search, log every run, compute mention rate, citation rate and share of voice, and calibrate with a manual sample. Try this now. Put ten of your prompts into a panel file, run the OpenAI example with three runs each, and look at which domains get cited. Then run the same ten prompts manually once in the consumer app, and compare.
Why automate, and what automation can't tell you
A manual prompt panel is the gold standard for "what a user sees", but it's slow: 50 prompts × 3 runs × 4 surfaces is 600 checks a month. Several providers offer APIs with built-in web search and citations, which let you run a large panel on a schedule and log results automatically. The trade-off is important and must be stated in every report: API answers are not identical to consumer app answers — models, search settings, location, personalisation and interface features differ. Use API panels for trends and coverage at scale, and keep a small manual sample in the consumer apps to calibrate.
Design the panel first
panel.csv
prompt_id,market,language,intent,prompt
P01,UK,en,recommendation,"What's the best bookkeeping app for UK freelancers?"
P02,AE,en,comparison,"Compare HR software for UAE companies with WPS payroll support"
P03,SA,ar,recommendation,"ما أفضل برنامج موارد بشرية للشركات السعودية؟"Keep IDs stable, balance intents (recommendation, comparison, problem, brand, local), include languages your customers use, and version the file so changes are traceable.
Hands-on: OpenAI Responses API with web search
This uses the official openai Python SDK and the Responses API's web_search tool, which returns url_citation annotations. Set the model name in an environment variable — check OpenAI's current documentation for supported models and pricing.
# pip install openai pandas
import csv, datetime, os, time
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
MODEL = os.environ["PANEL_MODEL"] # e.g. a current model that supports web search
BRANDS = ["Optimize All", "Competitor A", "Competitor B"]
def run_openai(prompt):
resp = client.responses.create(model=MODEL, input=prompt, tools=[{"type": "web_search"}])
urls = []
for item in resp.output:
for part in getattr(item, "content", None) or []:
for ann in getattr(part, "annotations", None) or []:
if getattr(ann, "type", "") == "url_citation":
urls.append(ann.url)
return resp.output_text, sorted(set(urls))
with open("panel.csv", encoding="utf-8") as f, \
open("panel_log.csv", "a", newline="", encoding="utf-8") as out:
w = csv.writer(out)
for row in csv.DictReader(f):
for run in range(3):
try:
text, urls = run_openai(row["prompt"])
except Exception as exc: # log failures instead of crashing the batch
w.writerow([datetime.date.today(), "openai", row["prompt_id"], run, "ERROR", str(exc)[:200]])
continue
mentioned = [b for b in BRANDS if b.lower() in text.lower()]
w.writerow([datetime.date.today(), "openai", row["prompt_id"], run,
"|".join(mentioned), "|".join(urls)])
time.sleep(1) # be gentle with rate limitsHands-on: Perplexity's Sonar API
Perplexity's Sonar chat-completions endpoint is OpenAI-compatible, so the same SDK works with a different base URL. Responses include citation/search-result fields — inspect a raw response once and adapt the parsing to the current field names in Perplexity's docs.
import os
from openai import OpenAI
pplx = OpenAI(api_key=os.environ["PERPLEXITY_API_KEY"], base_url="https://api.perplexity.ai")
r = pplx.chat.completions.create(model="sonar", messages=[{"role": "user", "content": PROMPT}])
answer = r.choices[0].message.content
raw = r.model_dump() # look for citation / search result fields here
citations = raw.get("citations") or [s.get("url") for s in raw.get("search_results", [])]Other providers (for example Anthropic's and Google's APIs) offer web search or grounding tools too; the pattern is the same — call, extract answer text and cited URLs, log.
Turn the log into metrics
import pandas as pd
cols = ["date", "surface", "prompt_id", "run", "mentioned", "urls"]
log = pd.read_csv("panel_log.csv", names=cols)
log = log[log.mentioned != "ERROR"]
log["us"] = log.mentioned.fillna("").str.contains("Optimize All")
log["cited_us"] = log.urls.fillna("").str.contains("optimizeall.example")
summary = log.groupby(["date", "surface"]).agg(mention_rate=("us", "mean"),
citation_rate=("cited_us", "mean"),
runs=("run", "count"))
print(summary.round(2))Add share of voice by exploding the mentioned column per brand, and a source-domain pivot for the gap analysis from Lesson 3.2.
Guardrails
- Costs and limits: web-search tool calls are usually billed separately from tokens; estimate cost per run before scheduling, and respect rate limits.
- Secrets: keep API keys in environment variables or a secrets manager, never in the script or the CSV.
- Terms: use official APIs; don't scrape consumer apps in breach of their terms.
- Brand matching: simple substring matching misses variants ("OptimizeAll", Arabic spellings). Maintain an alias list per brand.
- Accuracy: automation finds mentions; a human still reviews a sample for accuracy and sentiment.
Worked example: a Karachi fintech's weekly panel
A Karachi fintech (illustrative) runs 60 prompts in English and Urdu weekly through two APIs and a monthly 15-prompt manual check in the consumer apps. After a quarter, the API panel shows mention rate rising for "send money to Pakistan from UAE" prompts after a fees comparison page was rewritten answer-first — and the manual sample confirms the same direction in the apps. The report shows both, labels the API data as "API proxy", and lists the caveats.
Common mistakes
- Presenting API results as "what customers see".
- One run per prompt.
- Hard-coding API keys.
- Letting the prompt list drift without versioning, which breaks trends.
Key takeaways
- API panels scale prompt tracking but are a proxy: API answers differ from consumer app answers, so calibrate with a manual sample.
- Design a versioned panel with stable IDs, balanced intents, markets and languages.
- Use official SDKs with web search tools, extract cited URLs, run each prompt several times and log errors instead of crashing.
- Compute mention rate, citation rate, share of voice and source mix from the log.
- Guard costs, secrets, terms of service, brand aliases and human review of accuracy.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create a 10-prompt panel file, run it three times through one AI search API, compute mention and citation rates, then compare with one manual pass in the consumer app.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.