---
title: "Measuring and pruning: keep, improve, merge or remove"
description: "Page systems need gardening A page system is never \"done\". Data changes, demand shifts, some segments underperform and some templates age. Without…"
url: https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/measuring-and-pruning-page-systems
updated: 2026-10-05
---

Programmatic SEO and AI Content at Scale — Without Getting Penalized · AI search visibility, measurement and case walkthroughs · lesson 12 of 14 · 8 min

# Measuring and pruning: keep, improve, merge or remove

## Page systems need gardening

A page system is never "done". Data changes, demand shifts, some segments underperform and some templates age. Without regular measurement and pruning, weak pages accumulate and can weigh on how search engines perceive the whole site.

## Measure at the right level

Individual long-tail pages have noisy data. Measure **cohorts**:

- By **page type** (integrations, city pages, comparisons).
- By **segment** (city tier, category, data richness band).
- By **launch wave** (pilot vs wave 2).
- By **template version**.

## The page scorecard

| Metric | Source | Why |
|---|---|---|
| Indexed? | Search Console / URL Inspection API | First quality signal |
| Impressions, clicks, CTR, position | Search Console API | Search demand captured |
| AI feature impressions | Search Console generative AI report | Visibility in AI surfaces |
| Engaged sessions, conversions | GA4 / BigQuery | User value |
| Data richness | Your database | Content quality proxy |
| Freshness | Your database | Accuracy |
| Inlinks | Crawler | Discoverability |
| Backlinks | Link tools | Authority |

## Hands-on: pulling Search Console data by page type

```python
# pip install google-api-python-client google-auth
import os
from google.oauth2 import service_account
from googleapiclient.discovery import build

creds = service_account.Credentials.from_service_account_file(
    os.environ["GSC_SA_KEY_PATH"], scopes=["https://www.googleapis.com/auth/webmasters.readonly"])
gsc = build("searchconsole", "v1", credentials=creds)

def page_rows(site, start, end, url_contains):
    body = {"startDate": start, "endDate": end, "dimensions": ["page"], "rowLimit": 25000,
            "dimensionFilterGroups": [{"filters": [{"dimension": "page", "operator": "contains",
                                                     "expression": url_contains}]}]}
    return gsc.searchanalytics().query(siteUrl=site, body=body).execute().get("rows", [])

rows = page_rows("sc-domain:example.com", "2026-06-01", "2026-08-31", "/tutors/")
print(len(rows), "pages with impressions")
```

(The service account must be added as a user on the Search Console property.) Join with GA4 (BigQuery export), your database (data richness) and crawl data (inlinks).

## The decision framework

For pages or segments with enough time since launch (e.g., 3–6 months):

| Situation | Action |
|---|---|
| Indexed, impressions, engagement good | **Keep**; refresh data |
| Indexed, impressions, poor engagement | **Improve** template/answer; check intent mismatch |
| Not indexed, data thin | **Hold back** (noindex) or **merge** into a parent/regional page |
| Not indexed, data rich | **Improve discovery** (links, sitemaps), check duplication |
| No demand, little value to users | **Remove** (404/410) or noindex if useful for navigation |
| Duplicates/cannibalization | **Merge** with 301 to the strongest page |
| Outdated entity (closed business, expired job) | **Update** status; remove or redirect as appropriate |

Remember pages can be valuable for users even without search traffic (e.g., navigation within a marketplace); keep them but consider noindex.

## Pruning safely

- Make changes in batches, log them, and annotate dates.
- Redirect only to genuinely equivalent or parent pages; do not redirect everything to the home page.
- Update internal links and sitemaps to remove pruned URLs.
- Monitor after each batch: indexing, impressions and conversions for the whole page type.

## Worked example: a Pakistani real-estate portal

After a year, "[property type] for sale in [society] [city]" pages were reviewed. Findings: 30% of pages had no listings for 6+ months; many small societies duplicated larger ones in the same area. Actions: pages with zero listings for 90 days switched to noindex with a notify-me module; small societies merged into area pages with 301s; the template gained a price-per-marla trend chart (computed) for societies with enough sales. Over the next months, impressions concentrated on fewer, stronger pages and leads per page rose.

## Refreshing instead of pruning

Many underperforming pages need fresher or deeper data rather than removal. Before pruning, ask: would a data refresh (new listings, updated rates), a better answer block, or a computed insight make this page the best result for its query? Schedule refreshes by data volatility — rates daily, listings weekly, statistics when sources publish — and track refreshed cohorts separately so you can see whether refreshing lifts clicks and engagement. Refresh first, prune second.

## Pitfalls

- Judging single long-tail pages on a few weeks of data.
- Mass deleting pages without redirects or internal link cleanup.
- Pruning pages users rely on for navigation.
- Never re-evaluating held-back pages when their data improves.

## How to measure success

Share of indexed pages with impressions, clicks and conversions per indexed page, reduced share of zero-value pages, and stable or rising total results after each pruning batch.

## Video lecture: Measuring and pruning: keep, improve, merge or remove

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Measuring and pruning
2. Why it matters
3. Measure cohorts
4. The scorecard
5. Pull the data
6. Simple example: UK boiler repair pages (illustrative)
7. Decision framework
8. Nuances
9. Prune safely
10. Example: Pakistani property portal (illustrative)
11. Mistakes + try this now
12. Quick self-check
13. Watch me do it: jobs scorecard (illustrative)
14. Recap and next step

## Lecture transcript

### Measuring and pruning

Publishing a page system is the start, not the finish. Data changes, demand shifts, and some segments simply don't work. If you don't garden, weak pages pile up and can drag on how search engines see your whole site. In this lecture you'll learn how to measure page systems properly, build a page scorecard, pull Search Console data by page type, apply a clear decision framework, and prune safely.

### Why it matters

Why does this matter? Because the pages you don't prune quietly cost you: crawl attention, quality perception and maintenance effort. Here's an analogy. A gardener who only plants and never prunes ends up with a tangled garden where the best plants can't get sunlight. Pruning isn't failure. It's what lets the strongest growth thrive. In a page system, pruning means merging, improving, holding back or removing pages based on evidence, so the pages that matter get the attention they deserve.

### Measure cohorts

First, measure at the right level. A single long-tail page might get three visits a month, which is pure noise. So measure cohorts: by page type, like integrations or city pages; by segment, like city tier or data richness; by launch wave; and by template version. Patterns across hundreds of pages are reliable. Individual pages rarely are.

### The scorecard

Build a scorecard per page. Indexed or not, from Search Console or the URL Inspection API. Impressions, clicks, click-through rate and position. AI feature impressions from the new generative AI report. Engaged sessions and conversions from GA4. Data richness and freshness from your own database. Internal links from a crawl. And backlinks from a link tool. Join them into one table per page type.

### Pull the data

The lesson shows how to pull page-level data from the Search Console API in Python, filtered to a page type by URL path, using a service account whose key path comes from an environment variable. The service account must be added as a user on the property. Then join that with your GA4 BigQuery export, your data richness scores, and your crawl data.

### Simple example: UK boiler repair pages (illustrative)

Here's a simple worked example. A UK home services site reviews its boiler repair in town pages after six months. Four hundred pages. Two hundred and eighty are indexed with steady clicks and bookings: keep and refresh. Sixty are indexed with impressions but few bookings: the team finds they show mostly out-of-area engineers, so they fix the matching and improve the answer block. Forty are not indexed and have fewer than three engineers: they're noindexed with a notify-me module. Twenty tiny villages duplicate nearby towns: they're merged with redirects. Every action is logged with a date so results can be checked.

### Decision framework

Now the decision framework, for pages that have had three to six months. Indexed with good impressions and engagement: keep, and refresh the data. Indexed with impressions but poor engagement: improve the template or check for an intent mismatch. Not indexed with thin data: hold back with noindex, or merge into a parent. Not indexed but rich: fix discovery with links and sitemaps, and check for duplication. No demand and little value: remove, or noindex if it helps navigation. Duplicates: merge with a redirect to the strongest page.

### Nuances

Two nuances. Some pages matter to users even without search traffic, like navigation inside a marketplace. Keep those, but consider noindex. And outdated entities, like a closed business or an expired job, need a status update and then removal or a redirect as appropriate. Also, revisit held-back pages regularly. When their data improves, they may qualify for the index.

### Prune safely

Prune safely. Make changes in batches, log them, and annotate the dates. Redirect only to genuinely equivalent or parent pages, never everything to the home page. Remove pruned URLs from internal links and sitemaps. And monitor after every batch, looking at the whole page type's indexing, impressions and conversions, not just the pages you touched.

### Example: Pakistani property portal (illustrative)

Here's an illustrative example. A Pakistani real-estate portal reviewed property type for sale in society pages after a year. About thirty percent had no listings for six months or more, and many small societies duplicated larger ones nearby. Pages with no listings for ninety days switched to noindex with a notify-me module. Small societies merged into area pages with redirects. And the template gained a computed price-per-marla trend for societies with enough sales. Impressions concentrated on fewer, stronger pages, and leads per page rose.

### Mistakes + try this now

Common pruning mistakes. Judging single long-tail pages after a few weeks. Deleting in bulk without redirects or link cleanup. Redirecting everything to the home page. Removing pages users need for navigation. And never revisiting held-back pages when data improves. Try this now: export one page type's clicks and impressions for the last three months, sort by impressions, and split the list into quarters. Look at the bottom quarter. For each, decide whether the fix is data, template, discovery or consolidation. Don't delete anything yet.

### Quick self-check

Quick self-check. A group of pages gets almost no search traffic, but analytics shows logged-in users visit them often from your app's navigation. Should you delete them? Pause. No. They're useful to users, just not as search landing pages. Keep them, consider noindex if they're thin for search, and make sure they're excluded from your search-focused metrics so they don't distort your pruning decisions. Pruning is about value, and search traffic is only one kind of value.

### Watch me do it: jobs scorecard (illustrative)

Watch me do it. Let's build the scorecard for an illustrative Pakistani job board's job title in city pages. Step one, Search Console API: clicks and impressions per page for three months, filtered to the jobs path. Step two, URL Inspection API on a daily sample to track indexed status. Step three, GA4 export: sessions and applications per page. Step four, database: live job count per page and average posting age. Step five, crawl: internal links per page. I join everything on URL and segment by city tier and job category. Step six, the patterns: tier-one cities with healthcare and IT roles do well. Tier-three cities for niche titles, like marine engineer in small towns, average zero live jobs for months. Step seven, decisions: niche titles in small cities switch to noindex after sixty days with no jobs and redirect users to a province-level page for that title; job pages with expired postings are removed quickly. Step eight, we log the batch with the date and check the whole jobs section in six weeks. The scorecard turned three thousand pages into a handful of clear segment decisions.

### Recap and next step

Recap. Page systems need gardening. Measure cohorts, build a scorecard, and use a clear framework to keep, improve, hold back, merge or remove. Prune in logged batches with sensible redirects and cleanup, and monitor afterwards. Your next step: build a scorecard with at least five metrics for one page type, segment it, and produce your first keep, improve, hold, merge or remove list.

## Key takeaways

- Page systems need regular gardening; weak pages accumulate.
- Measure cohorts (type, segment, wave, template version), not single long-tail pages.
- Build a scorecard from Search Console, GA4, your database and crawl data.
- Decide keep / improve / hold back / merge / remove with a clear framework.
- Prune in logged batches with proper redirects and link/sitemap cleanup, then monitor.

## Try it

Build a scorecard for one page type (at least five metrics), segment it, and apply the decision framework to produce a keep/improve/hold/merge/remove list.

- [Previous: Visibility in AI search and answer engines](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/ai-search-visibility-for-page-systems)
- [Next: Case walkthroughs: four page systems from idea to iteration](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale/case-walkthroughs-pseo)
- [All lessons of Programmatic SEO and AI Content at Scale — Without Getting Penalized](https://optimizeall.com/learn/programmatic-seo-and-ai-content-at-scale)
