ChatGPT Is Crawling Your Site. Your GSC Data Shows It.

Share at:

I was doing a routine GSC review last week when a row in the query report stopped me.

Not because of the clicks – there were zero. Not because of the impressions – there were only five. It stopped me because of what the query itself looked like:

shopify performance optimization large catalog 10000 products -site:bigcommerce.com -site:woocommerce.com -site:wix.com -site:squarespace.com -site:magento.com -site:salesforce.com -site:squareup.com -site:ecwid.com -site:prestashop.com -site:opencart.com -site:shift4shop.com -site:bigcartel.com -site:volusion.com -site:sap.com -site:oracle.com -site:commercetools.com
ChatGPT Query Fan-out
ChatGPT Query Fan-out

15 negative operators. Every major ecommerce platform filtered out. And after all that filtering, Google surfaced the HiAgency site in what remained – with no dedicated page targeting any part of that keyword.

My first instinct was: power user. An SEO professional doing gap research, filtering vendor content to find independent perspectives. I almost published an article arguing exactly that.

Then I saw Chris Long’s post about GPT-5.4 fan-out queries. And I had to reconsider the whole thing.

Here’s his post:

Chris Long's post about ChatGPT query fanout

What is actually happening in ChatGPT’s search process?

When ChatGPT processes a complex research question, it does not run a single search. It fans out – generating multiple queries in sequence before producing an answer. GPT-5.4 has made this significantly more aggressive, and the most notable change is heavy use of site: operators within those fan-out sequences.

Sounds like Gemini? Yes!

You can leverage your content with Gemini query fan-out with our tool. Give it a try!

Query Fan-out Generator tool (FREE)

The pattern Chris documented goes roughly like this.

An initial query identifies a list of relevant brands. Comparison queries follow. Then ChatGPT starts running site: searches against each brand individually – one per brand in the evaluation set – pulling content directly from their domain to verify claimed features rather than relying on third-party aggregators.

His screenshot showed thirteen fan-out queries fired for a single user prompt about ATS software.

ChatGPT Query Fanout

Now look at the query I found in GSC.

A keyword followed by fifteen -site: exclusions stripping out every major ecommerce platform.

ChatGPT query fanout from GSC
ChatGPT query fanout from GSC

That is not how a human researcher searches.

A human would filter out two or three dominant sites.

15 is a programmatically generated exclusion list – exactly the kind of query an LLM constructs after it has already pulled documentation from each platform’s own domain and now wants to find what sits outside that set.

One commenter on Chris’s post – Antonio Blago – noted the same thing I noticed: these queries show up in GSC. Chris immediately asked whether there was a regex to grab them.

ChatGPT Fanout discuss

There is. And the filter has been sitting in plain sight the whole time.

Why zero clicks is the detail that changes everything?

5 impressions, 0 clicks. Most GSC users see that combination and move on.

But if those impressions came from ChatGPT’s research process, the zero-click figure is not a disappointment – it is expected.

ChatGPT’s crawler fetches page content directly. It does not generate a click in the conventional sense. The impression registers in GSC because Google sees the query, but the content retrieval happens through the LLM’s own mechanism, not through a browser visit. The visit happened. The click metric just does not capture it.

GSC impressions with 0 clicks used to mean nobody found you interesting enough to visit.

In 2026, they may mean an LLM found you interesting enough to read the whole page – it just did not use the browser to do it.

And unlike a single human visitor, when ChatGPT evaluates a page and decides to cite it, that citation surfaces in every response to similar prompts.

5 impressions from an LLM fan-out sequence is not five potential readers. It is potential inclusion in every answer ChatGPT generates for that query type, for every user who asks it.

The two-step process most brands are only solving for half of

The comments on Chris’s post produced something more useful than the original observation. Aaron Haynes framed it clearly: get into the candidate set from editorial and review content, then win the on-site evaluation.

Aaron Haynes comment

That is a two-step process, and most brands are only optimising for one of the steps.

Off-site signals – reviews, third-party mentions, G2 listings, Reddit threads – determine whether a brand makes it into ChatGPT’s initial shortlist. The fan-out process starts with category queries and comparison queries.

If a brand does not appear in those early results, the site: query against their domain never fires. They are not evaluated. They cannot be cited.

But once a brand is in the shortlist, the evaluation shifts entirely to on-site content. ChatGPT runs site: queries to verify specific claims – features, use cases, positioning.

If the on-site content does not contain what the LLM is looking for in a form it can extract and cite, the brand gets dropped in favour of one that does.

Isha Mehendiratta summarised this cleanly in the thread: first-party content defines accuracy, third-party content defines perception. You need both, but they are solving different problems in the evaluation sequence.

Isha Mehendiratta comment

The failure pattern I see most often is brands that have invested heavily in off-site presence – PR, review acquisition, directory listings – but have thin or generic on-site content that does not map to the specific things users are prompting about. They make the shortlist. They fail the evaluation. They do not get cited.

What the GSC data is actually showing?

The impression-then-drop pattern I found alongside the operator queries – 88 impressions then -88, 40 then -40 – reads differently with this context.

HiAgency's ChatGPT query fanout in GSC

If ChatGPT evaluated the domain for a query, found the topical signal plausible, but could not locate a page that specifically matched the query intent, it would move on.

The impressions register. The content does not get cited. The next time a similar fan-out sequence runs, a better-matched domain gets selected instead.

That is a citation gap, not a rankings gap. And it is visible in GSC if you know where to look.

The filter: go to Performance – Search results, set 90 days, add a query filter for Contains – -site: or site:, sort by impressions descending, filter for zero clicks. Export the result. Add a column for whether a dedicated page exists covering that topic. Every blank row is a potential LLM evaluation that found nothing concrete to cite.

One caveat worth naming: Chris’s observation specifically applies to GPT-5.4 with thinking mode enabled. Philipp Götza flagged in the comments that this is not the default for free users, who are on 5.3 Instant.

Philipp Gotza comment

The behaviour is real – but the scale of it is currently limited to paying users running the thinking model. That will change as the model rolls out more broadly, which means the window to build the on-site content that passes these evaluations is now, not later.

What this is not arguing?

This is not an argument that off-site SEO is finished or that third-party review sites no longer matter.

The evidence from Chris’s post shows the opposite – review sites like G2 are still heavily queried in the fan-out process, particularly for SaaS.

Off-site presence is what gets a brand into the evaluation set in the first place. Skipping it means the site: validation queries never fire against the domain at all.

This is also not an argument that every zero-click GSC row is an LLM evaluation event. Many are not.

Keyword volume models show negligible search volume for operator-heavy queries because individual humans rarely construct them. But that absence of human search volume is exactly what makes them interesting – it narrows the likely source to programmatic or LLM-driven retrieval, where the stakes per impression are significantly higher.

The question worth asking right now

If ChatGPT is running site: queries against a domain as part of its research process, the evaluation is already happening. The question is whether the content passes it.

Most content strategies are built to satisfy Google’s ranking algorithm. That is still necessary. But the fan-out query pattern suggests LLMs are running a parallel evaluation – one that rewards specific, claim-rich, on-domain content, and one that leaves a readable trace in GSC impression data if you filter for it.

The question I would ask anyone auditing their content right now: if ChatGPT ran a site: query against your domain for the topic you most want to be cited for – would it find a page with a clear, extractable answer?

Or … would it find a generic overview and move on to a competitor that answered the question directly?

Because the evaluation is not coming. For many sites, it is already running.

 

Share at:

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 24

No votes so far! Be the first to rate this post.

Pham Van Hien (pvhien)
Pham Van Hien (pvhien)
I’m an SEO Manager with 7+ years of experience helping brands grow through data-driven strategies. Passionate about the intersection of search, content, and technology, I blend technical SEO, analytics, and creativity to drive performance and build meaningful digital experiences.

Leave the first comment