AI Search Optimization: SEO for AI Overviews & Answer EnginesMeasuring AI visibility · Lesson 13 of 17

Automating a prompt panel with AI search APIs

Article · 16 min · 9 min lecture

Video lecture

Automating a prompt panel with AI search APIs

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Automating a prompt panel

  • APIs with web search and citations
  • A logged, scheduled pipeline
  • The caveat for every report

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why automate, and what automation can't tell you

A manual prompt panel is the gold standard for "what a user sees", but it's slow: 50 prompts × 3 runs × 4 surfaces is 600 checks a month. Several providers offer APIs with built-in web search and citations, which let you run a large panel on a schedule and log results automatically. The trade-off is important and must be stated in every report: API answers are not identical to consumer app answers — models, search settings, location, personalisation and interface features differ. Use API panels for trends and coverage at scale, and keep a small manual sample in the consumer apps to calibrate.

Design the panel first

panel.csv
prompt_id,market,language,intent,prompt
P01,UK,en,recommendation,"What's the best bookkeeping app for UK freelancers?"
P02,AE,en,comparison,"Compare HR software for UAE companies with WPS payroll support"
P03,SA,ar,recommendation,"ما أفضل برنامج موارد بشرية للشركات السعودية؟"

Keep IDs stable, balance intents (recommendation, comparison, problem, brand, local), include languages your customers use, and version the file so changes are traceable.

This uses the official openai Python SDK and the Responses API's web_search tool, which returns url_citation annotations. Set the model name in an environment variable — check OpenAI's current documentation for supported models and pricing.

# pip install openai pandas
import csv, datetime, os, time
from openai import OpenAI

client = OpenAI()                       # reads OPENAI_API_KEY from the environment
MODEL = os.environ["PANEL_MODEL"]       # e.g. a current model that supports web search
BRANDS = ["Optimize All", "Competitor A", "Competitor B"]

def run_openai(prompt):
    resp = client.responses.create(model=MODEL, input=prompt, tools=[{"type": "web_search"}])
    urls = []
    for item in resp.output:
        for part in getattr(item, "content", None) or []:
            for ann in getattr(part, "annotations", None) or []:
                if getattr(ann, "type", "") == "url_citation":
                    urls.append(ann.url)
    return resp.output_text, sorted(set(urls))

with open("panel.csv", encoding="utf-8") as f, \
     open("panel_log.csv", "a", newline="", encoding="utf-8") as out:
    w = csv.writer(out)
    for row in csv.DictReader(f):
        for run in range(3):
            try:
                text, urls = run_openai(row["prompt"])
            except Exception as exc:          # log failures instead of crashing the batch
                w.writerow([datetime.date.today(), "openai", row["prompt_id"], run, "ERROR", str(exc)[:200]])
                continue
            mentioned = [b for b in BRANDS if b.lower() in text.lower()]
            w.writerow([datetime.date.today(), "openai", row["prompt_id"], run,
                        "|".join(mentioned), "|".join(urls)])
            time.sleep(1)                     # be gentle with rate limits

Hands-on: Perplexity's Sonar API

Perplexity's Sonar chat-completions endpoint is OpenAI-compatible, so the same SDK works with a different base URL. Responses include citation/search-result fields — inspect a raw response once and adapt the parsing to the current field names in Perplexity's docs.

import os
from openai import OpenAI

pplx = OpenAI(api_key=os.environ["PERPLEXITY_API_KEY"], base_url="https://api.perplexity.ai")
r = pplx.chat.completions.create(model="sonar", messages=[{"role": "user", "content": PROMPT}])
answer = r.choices[0].message.content
raw = r.model_dump()                       # look for citation / search result fields here
citations = raw.get("citations") or [s.get("url") for s in raw.get("search_results", [])]

Other providers (for example Anthropic's and Google's APIs) offer web search or grounding tools too; the pattern is the same — call, extract answer text and cited URLs, log.

Turn the log into metrics

import pandas as pd
cols = ["date", "surface", "prompt_id", "run", "mentioned", "urls"]
log = pd.read_csv("panel_log.csv", names=cols)
log = log[log.mentioned != "ERROR"]
log["us"] = log.mentioned.fillna("").str.contains("Optimize All")
log["cited_us"] = log.urls.fillna("").str.contains("optimizeall.example")
summary = log.groupby(["date", "surface"]).agg(mention_rate=("us", "mean"),
                                               citation_rate=("cited_us", "mean"),
                                               runs=("run", "count"))
print(summary.round(2))

Add share of voice by exploding the mentioned column per brand, and a source-domain pivot for the gap analysis from Lesson 3.2.

Guardrails

  • Costs and limits: web-search tool calls are usually billed separately from tokens; estimate cost per run before scheduling, and respect rate limits.
  • Secrets: keep API keys in environment variables or a secrets manager, never in the script or the CSV.
  • Terms: use official APIs; don't scrape consumer apps in breach of their terms.
  • Brand matching: simple substring matching misses variants ("OptimizeAll", Arabic spellings). Maintain an alias list per brand.
  • Accuracy: automation finds mentions; a human still reviews a sample for accuracy and sentiment.

Worked example: a Karachi fintech's weekly panel

A Karachi fintech (illustrative) runs 60 prompts in English and Urdu weekly through two APIs and a monthly 15-prompt manual check in the consumer apps. After a quarter, the API panel shows mention rate rising for "send money to Pakistan from UAE" prompts after a fees comparison page was rewritten answer-first — and the manual sample confirms the same direction in the apps. The report shows both, labels the API data as "API proxy", and lists the caveats.

Common mistakes

  • Presenting API results as "what customers see".
  • One run per prompt.
  • Hard-coding API keys.
  • Letting the prompt list drift without versioning, which breaks trends.

Key takeaways

  • API panels scale prompt tracking but are a proxy: API answers differ from consumer app answers, so calibrate with a manual sample.
  • Design a versioned panel with stable IDs, balanced intents, markets and languages.
  • Use official SDKs with web search tools, extract cited URLs, run each prompt several times and log errors instead of crashing.
  • Compute mention rate, citation rate, share of voice and source mix from the log.
  • Guard costs, secrets, terms of service, brand aliases and human review of accuracy.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why must reports label API-based panel results as a proxy?
  2. Where should the API key live?
  3. A teammate rewrites 20 prompts mid-quarter without telling anyone. What's the main risk?

Put it into practice

Create a 10-prompt panel file, run it three times through one AI search API, compute mention and citation rates, then compare with one manual pass in the consumer app.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.