Quick answer: Since about September 13, Google’s anti-scraping measures have blocked most of the search-results data that third-party tools harvest — Nozzle’s Derek Perkins reported roughly an 80% drop in data he could pull, DataForSEO saw the same, and Sistrix flagged “reduced rate” collection on September 16. That’s a rank-tracker story on its face. The GEO angle almost nobody is asking: most “AI visibility” and “AI Share of Voice” dashboards scrape the very same Google results page to tell you whether you appear in AI Overviews and AI Mode. If the scrape rate just fell ~80%, those AI-citation numbers are now sampling a degraded, partial, possibly-stale feed — an already-inferred metric just got more inferred. The move: don’t panic-react to a dashboard dip that may be a data-collection artifact, keep the scope to Google surfaces only, and shift weight onto first-party signals you actually own — Search Console’s Generative AI report and your own GA4.
Trend watch — published September 20, 2026. This is a fact-check post: Google’s scraper crackdown escalated across the SEO community this week (Search Engine Roundtable recap, Sept 18; Nozzle and Sistrix status reports, Sept 13–16). We pulled the reported numbers, then asked the question the SEO coverage skipped — what does a data-collection outage do to AI visibility measurement — in the spirit of the GEO Lab.
What actually happened this week?
Around September 13, Google’s ability to block automated scraping of its search results took a visible step up. Derek Perkins of the rank tracker Nozzle reported that the majority of his tools’ attempts to grab data from Google Search were being blocked — roughly an 80% reduction in what he could collect — and that DataForSEO, a data supplier many other tools resell, saw the same hit. Sistrix’s status page noted on September 16 that “our data collection is currently running at a reduced rate,” and Glenn Gabe observed uneven results across vendors: Ahrefs scraping relatively effectively, Semrush slower, Sistrix well behind. Barry Schwartz rounded it up in Search Engine Roundtable’s September 18 recap.
This isn’t a single switch — it’s the latest turn in a year-long arc. In late July, Google began routing result links through a google.com/goto redirect that encodes the destination with Protocol Buffers, so a tool can’t decode the target URL locally and has to follow the full redirect through Google’s servers. Add the earlier removal of the num=100 parameter, the SerpApi litigation, and the scheduled end of the last official search API, and the direction is unambiguous: programmatic access to Google’s results is being tightened, deliberately and repeatedly. Vendors will engineer workarounds, as they always have; expect an ongoing disruption cycle rather than a clean permanent blackout.
Why should a GEO team care about a rank-tracker problem?
Because the plumbing is shared. When a dashboard tells you “you’re cited in 34% of AI Overviews for your keyword set,” it usually got that by scraping Google’s results page for each query and parsing which sources the AI Overview or AI Mode block linked to. That is the same scraping pipeline that just lost ~80% of its throughput. So the AI-visibility number you check on Monday is now built on a thinner, less reliable sample than the one you checked a month ago — even though the vendor’s chart still renders a confident line.
This is the quiet failure mode: a metric doesn’t announce when its underlying data source degrades. It just keeps drawing. A tool getting one-fifth of its usual Google responses may quietly fall back to cached results, sample fewer queries, or return gaps as apparent absences — any of which can look like “your AI citations dropped” when nothing about your actual visibility changed. As one rank-tracking write-up put it this week: be careful with the data from these tools right now, because it might be wrong.
Isn’t this just the “inferred data” problem you wrote about?
It’s that problem made physical. Two days ago we covered Duane Forrester’s warning that by 2027 you’ll be making decisions on inferred numbers you can’t actually check — a counted number comes from an event on infrastructure you or your vendor control; an inferred number is a sample extrapolated to a population nobody can fully enumerate. A scraped “AI Share of Voice” figure was always the inferred kind: a sample of queries, scraped from a surface you don’t own, extrapolated into a percentage. This week the sampling apparatus behind that inference lost 80% of its reach. The epistemology didn’t change; the sample just got worse, on a timeline Google controls and doesn’t announce.
That reframes the whole “AI visibility tracking” category. Any vendor number sourced from scraping Google is a rented signal — it exists at Google’s discretion and can be throttled without notice. The blended “AI Share of Voice” vanity score we’ve argued against was already the wrong metric for strategic reasons; now it has a supply-chain problem on top. Treat it as directional at best, and never as a KPI you’d defend to a board.
Does this mean your AI tracker is dead?
No — and this is where scope discipline matters. What broke is Google-surface scraping: AI Overviews and AI Mode citations that a tool reads off google.com. Tracking for ChatGPT, Perplexity, Claude, and Gemini’s app runs on a different pipeline — those tools query the assistants directly (via prompts or APIs), not by scraping Google’s results page — so a Google scraping crackdown doesn’t automatically dark out your non-Google measurement.
That distinction is the whole reason a single-source panic is a mistake. Our cross-engine testing found the four major engines shared only 2 of 395 cited domains, with about 85% unique to a single engine — which is exactly why there is no universal GEO strategy. A degradation confined to how you observe one engine’s surface tells you nothing about the other three, and doesn’t warrant tearing up your measurement model. Keep the blast radius honest: Google-surface tracking is impaired; per-engine coverage elsewhere is not.
Is your AI visibility actually dropping — or just your ability to see it?
These are different events, and this week they’re easy to confuse. Your citations are decided upstream, by how each engine retrieves and chooses sources; whether a third-party scraper can observe those citations is a separate, downstream question. A scraping outage moves the second without touching the first. So a dip in your AI-visibility dashboard right now is at least as likely to be a measurement artifact as a real change in how you’re cited.
We keep landing on this same distinction. Last week we argued that ranking volatility isn’t citation volatility, and that Search Console’s AI “position” isn’t a rank. Add this one to the set: a data-collection outage isn’t a visibility drop. The practical rule during a period like this — Google is also, separately, churning its ranking systems, with reversals around September 4 and 12 and an update on the 15th — is to freeze reactive content changes based on scraped dashboards. Don’t “fix” a number that may just be a broken feed.
What should you measure instead?
Shift weight from rented, scraped signals to counted ones you own — data that comes from infrastructure you or Google control, not from a scraper that can be switched off:
- Search Console’s Generative AI report — impressions and query coverage for AI Overviews and AI Mode come from Google’s own logs, not scraping, so this crackdown doesn’t touch it. Just remember its limits: impressions only, no click data, Google surfaces only.
- Your own analytics for the click and behavior side — our GA4 setup for AI traffic captures referred sessions and what they do next. It undercounts (referrer loss is real), but every session it records is an observed event, not an extrapolation.
- Server logs for crawler activity — GPTBot, ClaudeBot, and OAI-SearchBot hits are first-party facts about who’s fetching you, a useful complement to the citation picture and a reminder not to accidentally block the crawlers you want.
- Conversion value, not raw volume — a small AI-referred stream can still be your best one, as we found when asking whether AI search traffic converts better. Value survives even when volume counts get noisy.
- Lightweight per-engine presence checks — because Google-surface scraping is the impaired part, verify coverage engine by engine with the per-engine GEO checklist rather than trusting one blended score.
The practitioner takeaway
- Assume your scraped AI-visibility numbers are noisy right now. If a tool reads AI Overviews off Google, it lost ~80% of its data this week — the chart is confident; the sample isn’t.
- Don’t confuse a data outage with a visibility drop. Your citations are set upstream; a scraper losing access doesn’t change them, only your view of them.
- Keep the scope to Google. ChatGPT, Perplexity, Claude, and Gemini tracking run on separate pipelines — don’t let a Google crackdown trigger an all-engine panic.
- Freeze reactive changes. With Google also churning rankings this month, “fixing” a scraped number risks chasing an artifact.
- Move weight to counted, first-party signals. Search Console impressions, your own GA4, and server logs come from infrastructure you or Google control — they can’t be throttled out from under you.
Bottom line: Google didn’t take away your AI visibility this week — it made it harder for third parties to scrape the surface where part of that visibility is measured. If your dashboard’s number wobbles, the honest first hypothesis is a degraded feed, not a lost citation. The teams that already built their reporting on counted, first-party signals barely feel this. Everyone renting a scraped “AI Share of Voice” line just learned, again, that a metric you don’t own is a metric that can be switched off without asking you.
Frequently asked questions
What did Google change about scraping in September 2026?
Around September 13, Google’s anti-scraping measures became markedly more effective at blocking automated collection of its search results. Nozzle reported roughly an 80% drop in data it could scrape, DataForSEO saw the same, and Sistrix flagged reduced-rate collection on September 16. It builds on earlier moves — the google.com/goto redirect that hides destination URLs, removal of the num=100 parameter, and the winding down of official search API access.
Does this crackdown affect AI visibility tools too?
For Google surfaces, yes. Many AI visibility and “AI Share of Voice” dashboards scrape Google’s results page to detect whether you appear in AI Overviews and AI Mode, so they rely on the same pipeline that just lost most of its throughput. Their Google-surface citation numbers are now built on a thinner, less reliable sample. Tracking for ChatGPT, Perplexity, Claude, and Gemini uses separate pipelines and isn’t directly affected.
If my AI visibility number dropped this week, did I lose citations?
Not necessarily. Your citations are decided upstream by how each engine retrieves and selects sources; whether a scraper can observe them is a separate, downstream question. A scraping outage can lower a dashboard’s reading without any change in how you’re actually cited. During this period, treat a dip as a possible measurement artifact first, and confirm against first-party data before reacting.
What AI measurement isn’t affected by the scraping block?
Counted, first-party signals. Search Console’s Generative AI report draws impressions and query coverage from Google’s own logs rather than scraping, so it’s unaffected (though it still lacks click data). Your own GA4 captures AI-referred sessions and behavior, and server logs record AI crawler hits. These come from infrastructure you or Google control and can’t be throttled out from under you.
Will rank trackers and AI dashboards recover?
Vendors will likely engineer workarounds, as they have through previous changes, so expect partial recovery followed by more disruption rather than a clean permanent blackout. The structural lesson holds regardless: any metric sourced from scraping Google is a rented signal that can be throttled without notice, so it belongs alongside — not instead of — the first-party data you own.

Leave a Reply