BBA Index Coverage Audit

bestbrokersaustralia.org · GSC property sc-domain:bestbrokersaustralia.org · audited 2026-10-07 · for verification by a second agent
⚠️ The Phase-1 index is inverted — and there is a live leak feeding it.
Only 33 of 305 declared URLs are indexed, while ~4,100+ undeclared broker URLs are. And the true indexable surface is 20,818 broker URLs, of which only ~2,519 have surfaced so far — leaving an 18,299-URL crawl backlog still draining.
REVISION 2 — post-review correction. The first version understated the scale (I reported 4,176 indexed; Nick sees ~9,000 in GSC). It was right to be challenged:
305
URLs declared
33
…indexed
272
…not indexed
4,176
Indexed (floor, API)
~9,000
Indexed (GSC UI)
20,818
Indexable broker URLs
18,299
Crawl backlog pending
40
Clicks 3mo

1. Do not trust the supplied coverage export

The zip Nick supplied is unusable for this question. Table.csv holds 999 data rows with header URL,Last crawled. Google caps this export at 1,000 rows and sorts alphabetically, so it silently returned rows 1–999 and truncated everything else. There is no index-status column at all — it cannot answer "indexed or not" even if complete.

Consequence: 943 of the 999 rows are /broker/ URLs (clean + .html). Read naively it suggests "943 broker pages indexed", which is pure alphabetical-truncation artefact.

Chart.csv shows affected pages jumping 183 → 8,995 on 2026-09-19 and 8,995 → 12,373 on 2026-09-22 — that is the old bulk-sitemap crawl wave landing, and it created the current mess.

2. Declared vs actually indexed, by group

GroupIn sitemapIndexedState
/broker54097⚠ LEAKED — 4097 indexed, only 5 declared
/suburbs23028PARTIAL — 28 of 230 indexed
/mortgage-brokers016⚠ LEAKED — indexed, NOT declared
/state84OK
/self-employed73PARTIAL — 3 of 7 indexed
⚠ www. + param leaks03⚠ LEAKED — indexed, NOT declared
/calculator12OK
/asset-finance-brokers02⚠ LEAKED — indexed, NOT declared
/customs-brokers02⚠ LEAKED — indexed, NOT declared
/expat22OK
/insurance-brokers01⚠ LEAKED — indexed, NOT declared
/complaints11OK
/medico71PARTIAL — 1 of 7 indexed
/complaints.html01⚠ LEAKED — indexed, NOT declared
/matcher.html01⚠ LEAKED — indexed, NOT declared
/privacy11OK
/terms.html01⚠ LEAKED — indexed, NOT declared
/turnaround131PARTIAL — 1 of 13 indexed
/search.html01⚠ LEAKED — indexed, NOT declared
/disclaimer.html01⚠ LEAKED — indexed, NOT declared
/ (root page)11OK
/disclaimer11OK
/search11OK
/matcher11OK
/real-estate-agents01⚠ LEAKED — indexed, NOT declared
/terms11OK
/wealth-advisers01⚠ LEAKED — indexed, NOT declared
/construction20DECLARED — NOT INDEXED
/refinance40DECLARED — NOT INDEXED
/commercial30DECLARED — NOT INDEXED
/bad-credit10DECLARED — NOT INDEXED
/about10DECLARED — NOT INDEXED
/valuation20DECLARED — NOT INDEXED
/contract20DECLARED — NOT INDEXED
/auction30DECLARED — NOT INDEXED
/first-home-buyer70DECLARED — NOT INDEXED

Index rate per sitemap

SitemapIndexedTotalRate
sitemap-pages.xml111764%
sitemap-suburbs.xml152306%
sitemap-scenarios.xml75313%
sitemap-brokers-verified.xml050%

All 5 URLs in sitemap-brokers-verified.xml are unindexed — the cleanest single proof of the inversion.

2A. 🚨 The leak — how 10,409 files became 20,818 indexable URLs

This is the single most important finding. The problem is not a stale index to be cleaned up. It is a live leak: as long as unknown /broker/* paths are served out of a 10,409-file bucket with index, follow, Google keeps discovering new dossiers indefinitely. Deindex 5,000 today and 5,000 more take their place.

Tier-2 bucket inventory — gs://bestbrokersaustralia-static/

PrefixObjects
broker/10,409
assets/238
suburbs/235
calculator/231
turnaround/13
state/9
first-home-buyer/, medico/, self-employed/21
others52
TOTAL11,208

All 10,409 broker/ objects are .html.

The local tree is blind to 98.7% of them

Source of truthBroker files
v3psycho/broker/*.html (local tree, CI, sync_core.py)132
gs://bestbrokersaustralia-static/broker/10,409
GCS-only — invisible to the local tree10,277

Empirical proof: 120 of 120 randomly sampled indexed broker URLs were absent from the local tree, yet 63% returned 200 + index, follow on the live site. They are served from GCS.

The mechanism — two lines in docs-core/_worker.js

// line ~259-262 — extensionless -> .html -> GCS
if (!pathname.includes('.') && pathname !== '/') {
  pathname += '.html';
}

// line ~181 — .html -> clean 301, with NO X-Robots-Tag
if (url.pathname.endsWith('.html') && ...) {
  return Response.redirect(url.origin + cleanPath + url.search, 301);
}

// line ~160-178 — noindex covers only 8 hardcoded routes
const NOINDEX_INTERNAL_ROUTES = new Set([...]);   // /broker/* NOT included
const applyNoindexHeader = (pathname, res) => {
  if (!NOINDEX_INTERNAL_ROUTES.has(pathname)) return res;   // falls straight through
};

So every GCS file is indexable two ways:

URL formBehaviour
/broker/<slug>appends .html → fetches from GCS → 200 + index, follow
/broker/<slug>.htmlbare 301, no x-robots-tag → indexed as a separate URL

10,409 × 2 = 20,818 indexable broker URLs.

Leak progress

StageCount
Indexable broker URLs at edge20,818
Surfaced in Google — clean1,441
Surfaced in Google — .html1,078
Pending crawl backlog18,299

Chart.csv records affected pages climbing 183 → 8,995 → 12,373 between 2026-09-19 and 2026-09-22. That curve has not plateaued — it is this backlog draining.

3. What is indexed but should NOT be

Leaked groupIndexed
/mortgage-brokers16
/suburbs13
⚠ www. + param leaks3
/asset-finance-brokers2
/customs-brokers2
/calculator1
/complaints.html1
/disclaimer.html1
/insurance-brokers1
/matcher.html1
/real-estate-agents1
/search.html1
/state1
/terms.html1
/wealth-advisers1

Plus 4,097 /broker/ dossier URLs, where only 5 are declared. The legacy satellite dirs above are absent from the local tree and return 404, yet remain in the index.

Declared groups still not indexed (excluding /suburbs)

GroupNot indexed
/turnaround12
/first-home-buyer7
/medico6
/state5
/broker5
/refinance4
/self-employed4
/auction3
/commercial3
/construction2
/contract2
/valuation2
/about1
/bad-credit1

4. How much of /broker/ is dead?

Random sample, n=120 indexed broker URLs (curl -L, follow redirects):

ResultCount
76200
44404

So 37% (95% CI 29–46%) are dead 404 pages. Projected across 4,097 indexed broker URLs: ~1,502 dead 404s holding index slots and ~2,595 live dossier pages that should be noindexed.

They stay indexed because the 404 page emits noindex, nofollow — but a 404 only deindexes on recrawl, and thousands of URLs are ahead of them in the queue.

Duplicate structure: 4,097 broker URLs → 3,323 distinct slugs → 774 slugs indexed twice (clean + .html), 961 .html-only stragglers, 1,735 .html URLs total.

5. URL Inspection sample (n=50)

CountCoverage state
19Discovered – currently not indexed
17Submitted and indexed
12URL is unknown to Google
2Duplicate, Google chose different canonical than user

6. Secondary defects

Canonical mismatches in sitemap

URLCanonical points to
/broker/beat-my-home-loan-sydney/broker/david-chi-tran
/broker/emerge-finance-ashgrove/broker/emerge-finance

The mechanical cause of the duplicate bloat

$ curl -sI https://bestbrokersaustralia.org/broker/ryker-capital-ingleburn.html
HTTP/2 301
location: https://bestbrokersaustralia.org/broker/ryker-capital-ingleburn
# <- no x-robots-tag header

A bare 301 does not deindex the source URL. Google keeps the .html URL as a separate index entry indefinitely. This one omission accounts for the 1,735 duplicate broker URLs plus 5 root-level .html twins.

7. Root cause

Two competing indexes exist at once: an old index of ~3,300 broker dossiers plus 6 legacy satellite directories from the bulk-sitemap era, and the Phase-1 index where only 33 of 305 have landed. Google burns crawl budget on 404s and .html duplicates while 272 Phase-1 URLs sit in "Discovered – currently not indexed". On a domain where ~37% of URLs are dead and ~42% are duplicates, overall site-quality assessment is dragged down and the good pages are suppressed.

The sitemaps were never the problem. The index was never cleaned.

8. Remediation plan — REVISED (stop the leak first)

Step 0 is mandatory and comes before everything else. The original plan treated a fixed stale index; §2A proves the supply is still live. Until the worker stops minting indexable broker URLs, every deindexing effort gets refilled.
#ActionWhere
0aAdd X-Robots-Tag: noindex to the .html → clean 301_worker.js ~181
0bAdd X-Robots-Tag: noindex to Tier-2 GCS responses for /broker/* (needs a prefix rule — the allowlist at line 160 misses it)_worker.js ~307
0cDecide the fate of the 10,409 GCS dossiers — prune the broker/ prefix, or move it behind an unlinked pathGCS
1410 (not 404) all dead dossier slugs — deindexes ~2× faster, stops recrawl churnscripted
2410 the satellite footprints (/mortgage-brokers/, /asset-finance-brokers/, /customs-brokers/, /insurance-brokers/, /real-estate-agents/, /wealth-advisers/)_redirects
3Fix /privacy + /terms canonical conflicts2 files
4Fix 2 cross-broker canonicals2 files
5Block ?q= leak; canonicalise www. → apexconfig
6Dedupe 5 <title> sets; fix /suburbs/index 308small
7Only then: rebuild internal linking into unindexed Phase-1 groups (unlocks 272 URLs)content

Verify step 0 landed: curl -sI https://bestbrokersaustralia.org/broker/<any-slug> and confirm x-robots-tag: noindex appears.

Sequencing matters: adding content (step 7) while the worker serves 20,818 indexable broker URLs just feeds more pages into a domain Google is already discounting.

8b. Original table (superseded, kept for reference)

#ActionKills
1Add X-Robots-Tag: noindex to every .html → clean 301 in v3psycho/_redirects~1,740 dupes
2410 (not 404) all dead dossier slugs — deindexes ~2× faster, stops recrawl churn~1,500 dead
3410 the satellite footprint groups (/mortgage-brokers/, /asset-finance-brokers/, /customs-brokers/, /insurance-brokers/, /real-estate-agents/, /wealth-advisers/)23 legacy
4Fix /privacy + /terms canonical conflicts2
5Fix 2 cross-broker canonicals2
6Rebuild internal linking into unindexed Phase-1 groups (/first-home-buyer 0/7, /refinance 0/4, /auction 0/3, /turnaround 1/13, /suburbs 28/230)unlock 272
7Block ?q= leak; canonicalise www. → apex3
8Dedupe 5 <title> sets; fix /suburbs/index 3086
Sequencing matters: steps 1–3 first. Adding content (step 6) before removing the dead 4,143 URLs just adds pages to a domain Google is already discounting.

9. Independent verification checklist

  1. Highest priority: confirm gs://bestbrokersaustralia-static/ holds 10,409 objects under broker/ while v3psycho/broker/ holds 132. Read-only devstorage.read_only scope; script in audit/gcs_evidence.md.
  2. Confirm _worker.js line ~260 appends .html to extensionless /broker/* and serves it from GCS — each file indexable under two URLs.
  3. Confirm applyNoindexHeader (line ~169) covers only 8 hardcoded routes, so /broker/* emits no x-robots-tag.
  4. If you have GSC UI access, screenshot the Pages report breakdown — that is the one number this audit cannot verify independently, and the most valuable missing input. Treat ~9,000 as authoritative for "indexed"; 4,176 is a reproducible API floor.
  5. Confirm sitemap-brokers-verified.xml holds exactly 5 URLs, 0 indexed.
  6. Confirm /suburbs declares 230 but only ~28 appear in Google (6%).
  7. Confirm /mortgage-brokers/barton/ returns 404 yet still draws GSC impressions — proves stale-index retention.
  8. Confirm a 301 from any .html URL carries no x-robots-tag — the mechanical cause of the duplicate bloat.
  9. Re-run the GSC Search Analytics query below and confirm ≥4,000 /broker/ URLs are returned.
from google.oauth2 import service_account
from googleapiclient.discovery import build
creds = service_account.Credentials.from_service_account_file(
    'service_account.json',
    scopes=['https://www.googleapis.com/auth/webmasters.readonly'])
sc = build('searchconsole', 'v1', credentials=creds)
sc.searchanalytics().query(
    siteUrl='sc-domain:bestbrokersaustralia.org',
    body={'startDate': '2026-06-01', 'endDate': '2026-10-07',
          'dimensions': ['page'], 'rowLimit': 25000, 'dataState': 'all'}
).execute()

10. Methodology caveats — read before disputing

Generated by DSH agent (Sick) · 2026-10-07 · source: audit/INDEX_COVERAGE_AUDIT_2026-10-07.md · evidence: audit/evidence_bundle.json