sc-domain:bestbrokersaustralia.org · audited 2026-10-07 · for verification by a second agentrowLimit=25000 and across a 16-month window), so the query was correct but the metric was a floor. ~2,006 of those pages had just one impression; indexed-but-never-shown pages are invisible to Search Analytics.gs://bestbrokersaustralia-static/ holds 11,208 objects, 10,409 under broker/. The local v3psycho/ tree knows only 132. The problem is not a stale index to clean — it is a leaking bucket. See §4A.Table.csv was alphabetically sorted. False — verified rows == sorted(rows) is False both ways. The export is still truncated and unusable, but that specific claim was an unverified assertion.Table.csv holds 999 data rows with header URL,Last crawled. Google caps this export at 1,000 rows and sorts alphabetically, so it silently returned rows 1–999 and truncated everything else. There is no index-status column at all — it cannot answer "indexed or not" even if complete.
Consequence: 943 of the 999 rows are /broker/ URLs (clean + .html). Read naively it suggests "943 broker pages indexed", which is pure alphabetical-truncation artefact.
Chart.csv shows affected pages jumping 183 → 8,995 on 2026-09-19 and 8,995 → 12,373 on 2026-09-22 — that is the old bulk-sitemap crawl wave landing, and it created the current mess.
| Group | In sitemap | Indexed | State |
|---|---|---|---|
/broker | 5 | 4097 | ⚠ LEAKED — 4097 indexed, only 5 declared |
/suburbs | 230 | 28 | PARTIAL — 28 of 230 indexed |
/mortgage-brokers | 0 | 16 | ⚠ LEAKED — indexed, NOT declared |
/state | 8 | 4 | OK |
/self-employed | 7 | 3 | PARTIAL — 3 of 7 indexed |
⚠ www. + param leaks | 0 | 3 | ⚠ LEAKED — indexed, NOT declared |
/calculator | 1 | 2 | OK |
/asset-finance-brokers | 0 | 2 | ⚠ LEAKED — indexed, NOT declared |
/customs-brokers | 0 | 2 | ⚠ LEAKED — indexed, NOT declared |
/expat | 2 | 2 | OK |
/insurance-brokers | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/complaints | 1 | 1 | OK |
/medico | 7 | 1 | PARTIAL — 1 of 7 indexed |
/complaints.html | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/matcher.html | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/privacy | 1 | 1 | OK |
/terms.html | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/turnaround | 13 | 1 | PARTIAL — 1 of 13 indexed |
/search.html | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/disclaimer.html | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/ (root page) | 1 | 1 | OK |
/disclaimer | 1 | 1 | OK |
/search | 1 | 1 | OK |
/matcher | 1 | 1 | OK |
/real-estate-agents | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/terms | 1 | 1 | OK |
/wealth-advisers | 0 | 1 | ⚠ LEAKED — indexed, NOT declared |
/construction | 2 | 0 | DECLARED — NOT INDEXED |
/refinance | 4 | 0 | DECLARED — NOT INDEXED |
/commercial | 3 | 0 | DECLARED — NOT INDEXED |
/bad-credit | 1 | 0 | DECLARED — NOT INDEXED |
/about | 1 | 0 | DECLARED — NOT INDEXED |
/valuation | 2 | 0 | DECLARED — NOT INDEXED |
/contract | 2 | 0 | DECLARED — NOT INDEXED |
/auction | 3 | 0 | DECLARED — NOT INDEXED |
/first-home-buyer | 7 | 0 | DECLARED — NOT INDEXED |
| Sitemap | Indexed | Total | Rate |
|---|---|---|---|
sitemap-pages.xml | 11 | 17 | 64% |
sitemap-suburbs.xml | 15 | 230 | 6% |
sitemap-scenarios.xml | 7 | 53 | 13% |
sitemap-brokers-verified.xml | 0 | 5 | 0% |
All 5 URLs in sitemap-brokers-verified.xml are unindexed — the cleanest single proof of the inversion.
/broker/* paths are served out of a 10,409-file bucket with index, follow, Google keeps discovering new dossiers indefinitely. Deindex 5,000 today and 5,000 more take their place.gs://bestbrokersaustralia-static/| Prefix | Objects |
|---|---|
broker/ | 10,409 |
assets/ | 238 |
suburbs/ | 235 |
calculator/ | 231 |
turnaround/ | 13 |
state/ | 9 |
first-home-buyer/, medico/, self-employed/ | 21 |
| others | 52 |
| TOTAL | 11,208 |
All 10,409 broker/ objects are .html.
| Source of truth | Broker files |
|---|---|
v3psycho/broker/*.html (local tree, CI, sync_core.py) | 132 |
gs://bestbrokersaustralia-static/broker/ | 10,409 |
| GCS-only — invisible to the local tree | 10,277 |
Empirical proof: 120 of 120 randomly sampled indexed broker URLs were absent from the local tree, yet 63% returned 200 + index, follow on the live site. They are served from GCS.
docs-core/_worker.js// line ~259-262 — extensionless -> .html -> GCS
if (!pathname.includes('.') && pathname !== '/') {
pathname += '.html';
}
// line ~181 — .html -> clean 301, with NO X-Robots-Tag
if (url.pathname.endsWith('.html') && ...) {
return Response.redirect(url.origin + cleanPath + url.search, 301);
}
// line ~160-178 — noindex covers only 8 hardcoded routes
const NOINDEX_INTERNAL_ROUTES = new Set([...]); // /broker/* NOT included
const applyNoindexHeader = (pathname, res) => {
if (!NOINDEX_INTERNAL_ROUTES.has(pathname)) return res; // falls straight through
};
So every GCS file is indexable two ways:
| URL form | Behaviour |
|---|---|
/broker/<slug> | appends .html → fetches from GCS → 200 + index, follow |
/broker/<slug>.html | bare 301, no x-robots-tag → indexed as a separate URL |
10,409 × 2 = 20,818 indexable broker URLs.
| Stage | Count |
|---|---|
| Indexable broker URLs at edge | 20,818 |
| Surfaced in Google — clean | 1,441 |
Surfaced in Google — .html | 1,078 |
| Pending crawl backlog | 18,299 |
Chart.csv records affected pages climbing 183 → 8,995 → 12,373 between 2026-09-19 and 2026-09-22. That curve has not plateaued — it is this backlog draining.
| Leaked group | Indexed |
|---|---|
/mortgage-brokers | 16 |
/suburbs | 13 |
⚠ www. + param leaks | 3 |
/asset-finance-brokers | 2 |
/customs-brokers | 2 |
/calculator | 1 |
/complaints.html | 1 |
/disclaimer.html | 1 |
/insurance-brokers | 1 |
/matcher.html | 1 |
/real-estate-agents | 1 |
/search.html | 1 |
/state | 1 |
/terms.html | 1 |
/wealth-advisers | 1 |
Plus 4,097 /broker/ dossier URLs, where only 5 are declared. The legacy satellite dirs above are absent from the local tree and return 404, yet remain in the index.
| Group | Not indexed |
|---|---|
/turnaround | 12 |
/first-home-buyer | 7 |
/medico | 6 |
/state | 5 |
/broker | 5 |
/refinance | 4 |
/self-employed | 4 |
/auction | 3 |
/commercial | 3 |
/construction | 2 |
/contract | 2 |
/valuation | 2 |
/about | 1 |
/bad-credit | 1 |
Random sample, n=120 indexed broker URLs (curl -L, follow redirects):
| Result | Count |
|---|---|
| 76 | 200 |
| 44 | 404 |
So 37% (95% CI 29–46%) are dead 404 pages. Projected across 4,097 indexed broker URLs: ~1,502 dead 404s holding index slots and ~2,595 live dossier pages that should be noindexed.
noindex, nofollow — but a 404 only deindexes on recrawl, and thousands of URLs are ahead of them in the queue.Duplicate structure: 4,097 broker URLs → 3,323 distinct slugs → 774 slugs indexed twice (clean + .html), 961 .html-only stragglers, 1,735 .html URLs total.
| Count | Coverage state |
|---|---|
| 19 | Discovered – currently not indexed |
| 17 | Submitted and indexed |
| 12 | URL is unknown to Google |
| 2 | Duplicate, Google chose different canonical than user |
| URL | Canonical points to |
|---|---|
/broker/beat-my-home-loan-sydney | /broker/david-chi-tran |
/broker/emerge-finance-ashgrove | /broker/emerge-finance |
$ curl -sI https://bestbrokersaustralia.org/broker/ryker-capital-ingleburn.html HTTP/2 301 location: https://bestbrokersaustralia.org/broker/ryker-capital-ingleburn # <- no x-robots-tag header
A bare 301 does not deindex the source URL. Google keeps the .html URL as a separate index entry indefinitely. This one omission accounts for the 1,735 duplicate broker URLs plus 5 root-level .html twins.
<title> sets (5): Kingston ×3, Brighton ×2, Manly ×2, Paddington ×2, one Brisbane broker ×2./suburbs/index returns 308 — the only non-200 among the 305 declared URLs.noindex, and 304/305 return 200. The 272 unindexed pages are not blocked — Google is declining to index them. This rules out a robots/meta mistake and confirms a crawl-priority + site-quality problem./?q=Hobart; www. host duplicates.Two competing indexes exist at once: an old index of ~3,300 broker dossiers plus 6 legacy satellite directories from the bulk-sitemap era, and the Phase-1 index where only 33 of 305 have landed. Google burns crawl budget on 404s and .html duplicates while 272 Phase-1 URLs sit in "Discovered – currently not indexed". On a domain where ~37% of URLs are dead and ~42% are duplicates, overall site-quality assessment is dragged down and the good pages are suppressed.
| # | Action | Where |
|---|---|---|
| 0a | Add X-Robots-Tag: noindex to the .html → clean 301 | _worker.js ~181 |
| 0b | Add X-Robots-Tag: noindex to Tier-2 GCS responses for /broker/* (needs a prefix rule — the allowlist at line 160 misses it) | _worker.js ~307 |
| 0c | Decide the fate of the 10,409 GCS dossiers — prune the broker/ prefix, or move it behind an unlinked path | GCS |
| 1 | 410 (not 404) all dead dossier slugs — deindexes ~2× faster, stops recrawl churn | scripted |
| 2 | 410 the satellite footprints (/mortgage-brokers/, /asset-finance-brokers/, /customs-brokers/, /insurance-brokers/, /real-estate-agents/, /wealth-advisers/) | _redirects |
| 3 | Fix /privacy + /terms canonical conflicts | 2 files |
| 4 | Fix 2 cross-broker canonicals | 2 files |
| 5 | Block ?q= leak; canonicalise www. → apex | config |
| 6 | Dedupe 5 <title> sets; fix /suburbs/index 308 | small |
| 7 | Only then: rebuild internal linking into unindexed Phase-1 groups (unlocks 272 URLs) | content |
Verify step 0 landed: curl -sI https://bestbrokersaustralia.org/broker/<any-slug> and confirm x-robots-tag: noindex appears.
| # | Action | Kills |
|---|---|---|
| 1 | Add X-Robots-Tag: noindex to every .html → clean 301 in v3psycho/_redirects | ~1,740 dupes |
| 2 | 410 (not 404) all dead dossier slugs — deindexes ~2× faster, stops recrawl churn | ~1,500 dead |
| 3 | 410 the satellite footprint groups (/mortgage-brokers/, /asset-finance-brokers/, /customs-brokers/, /insurance-brokers/, /real-estate-agents/, /wealth-advisers/) | 23 legacy |
| 4 | Fix /privacy + /terms canonical conflicts | 2 |
| 5 | Fix 2 cross-broker canonicals | 2 |
| 6 | Rebuild internal linking into unindexed Phase-1 groups (/first-home-buyer 0/7, /refinance 0/4, /auction 0/3, /turnaround 1/13, /suburbs 28/230) | unlock 272 |
| 7 | Block ?q= leak; canonicalise www. → apex | 3 |
| 8 | Dedupe 5 <title> sets; fix /suburbs/index 308 | 6 |
gs://bestbrokersaustralia-static/ holds 10,409 objects under broker/ while v3psycho/broker/ holds 132. Read-only devstorage.read_only scope; script in audit/gcs_evidence.md._worker.js line ~260 appends .html to extensionless /broker/* and serves it from GCS — each file indexable under two URLs.applyNoindexHeader (line ~169) covers only 8 hardcoded routes, so /broker/* emits no x-robots-tag.sitemap-brokers-verified.xml holds exactly 5 URLs, 0 indexed./suburbs declares 230 but only ~28 appear in Google (6%)./mortgage-brokers/barton/ returns 404 yet still draws GSC impressions — proves stale-index retention..html URL carries no x-robots-tag — the mechanical cause of the duplicate bloat./broker/ URLs are returned.from google.oauth2 import service_account
from googleapiclient.discovery import build
creds = service_account.Credentials.from_service_account_file(
'service_account.json',
scopes=['https://www.googleapis.com/auth/webmasters.readonly'])
sc = build('searchconsole', 'v1', credentials=creds)
sc.searchanalytics().query(
siteUrl='sc-domain:bestbrokersaustralia.org',
body={'startDate': '2026-06-01', 'endDate': '2026-10-07',
'dimensions': ['page'], 'rowLimit': 25000, 'dataState': 'all'}
).execute()
random.seed(11)).audit/evidence_bundle.json (319 KB).audit/INDEX_COVERAGE_AUDIT_2026-10-07.md · evidence: audit/evidence_bundle.json