Skip to main content
Private betaThe Shopify app that makes your products readable and verifiable by AI.Shopify app: be readable by AII want in
AI Commerce

AI Traffic: The Real Crawl-to-Click Ratio

7 min read Updated Recently updated
#ai-commerce #ai-traffic #measurement #ga4 #crawlers #chatgpt #shopify
Share

For a few weeks we did something most stores never do: we logged every single AI bot request to our own site at the network edge, where JavaScript cannot hide anything, and compared it to what Google Search Console and GA4 reported for the same period. The gap between what the AI layer does to a site and what a store owner can see in their analytics turned out to be the whole story.

This is original research on a single, low-traffic site (our own), so read the numbers as an order of magnitude, not a benchmark. What makes them worth publishing is that every first-party figure below lines up with independent 2026 data from Cloudflare Radar, Similarweb and StatCounter. Method and honest limits are in their own section; nothing here is a modeled estimate dressed up as a measurement.

In 60 words

AI engines crawl constantly and refer almost nobody. On our edge logs, AI engines fetched pages 385 to 627 times a day (ChatGPT’s crawler was 74 to 79% of it, zero blocked), while the human-referral-to-fetch ratio was about 0.004. GA4 misses most of this: crawlers run no JavaScript, and AI-referred humans arrive with no referrer, landing in Direct.

AI traffic funnel: AI engines crawl a store 385 to 627 times a day while referring almost no human visitors, with Cloudflare Radar crawl-to-referral ratios (Google 5 to 1, Claude 23951 to 1) and the GA4 blind spot.
Figure 1: The AI layer crawls constantly and refers almost no one. Most of our own first-party funnel is invisible to GA4.

The paradox: massively crawled, almost never clicked

On a normal day, AI engines requested pages from our site between 385 and 627 times (counting only recognized AI engines, separated from other bots like ByteDance or CommonCrawl). ChatGPT’s crawler alone accounted for 74 to 79% of those recognized AI hits. Across every sampled day, zero AI crawlers were blocked: robots.txt and the WAF let them all through.

Against that: roughly 5 to 10 human sessions a day. The AI layer treats the site as something to read constantly and something to send visitors to almost never. That imbalance is the “crawl-to-click gap,” and it is not a quirk of our site. It is the structural shape of AI traffic in 2026.

How we measured it (and where it is thin)

Transparency is the argument here, so here are the four instruments and their limits:

  1. Cloudflare edge logs capture every request by user agent and status, including bots that never run JavaScript. This is the ground truth for crawl volume. Limit: the free plan caps each window at about 23 hours, so there is no native rolling 30-day view, just a stack of 23-hour days. Our series started June 26, 2026, so this is about 17 real days, not a full month.
  2. A server-side AI traffic monitor distinguishes an AI fetching a page (as a source) from a human arriving from an AI. Limit: the referral counts are point-in-time snapshots.
  3. Google Search Console for search impressions and clicks by market. This one is solid.
  4. GA4 for sessions and conversions. Limit: very small N, and polluted by datacenter spam bots (we saw impossible conversion counts from Nepal and Singapore), so GA4 conversion rates here are indicative at best.

Single site, micro-traffic, short window. Every claim below is framed accordingly.

What GA4 cannot see, and why

The first blind spot is structural. Most AI crawlers do not render JavaScript, and GA4 counts visits through a JavaScript tag. So the hundreds of daily AI fetches we see at the edge are, by construction, invisible to GA4. Any analytics built on client-side tags will under-report the AI crawl to zero, because the crawler never runs the tag.

The second blind spot is attribution. When a real human clicks through from an AI, the referrer is often missing, especially from mobile ChatGPT apps that strip it. Those visitors land in GA4’s “Direct / none” bucket, indistinguishable from someone typing your URL. An analysis of 446,405 visits (The Digital Bloom, February 2026) measured 70.6% of AI visits arriving with no referrer, and Averi’s 2026 analysis puts 30 to 50% of the AI pipeline hidden inside Direct (a related but distinct measure: misattributed to Direct rather than strictly no-referrer). We saw the same shape: a large Direct bucket that a naive read would attribute to brand awareness.

A note on rigor: we do not publish a single “GA4 undercounts by X times” multiplier, because we could not derive one cleanly from our data. The honest, sourced statement is the range above, from third parties: a large fraction of AI-referred humans is misattributed to Direct.

The crawl-to-refer gap, in numbers

How lopsided is the funnel? On our site, the ratio of human AI referrals to AI fetches (call it an “AI citation CTR”) was about 0.004: roughly one human referral for every 270 AI fetches, measured as 8 referrals against 2,166 live AI fetches over 14 days on our server-side monitor (the origin fetches, a subset of the higher edge crawl volume above).

That is a tiny sample, but it sits squarely inside the third-party ranges. Using Cloudflare Radar data, SEOmator (March 2026) reported crawl-to-referral ratios of roughly 23,951:1 for Anthropic/Claude, about 1,276:1 for OpenAI, around 111:1 for Perplexity, versus about 5:1 for Google. Cloudflare’s own foundational post (August 2025) framed the same phenomenon and noted that training accounts for roughly 80% of AI bot activity, search about 18%, and actual user-triggered actions about 2%. In other words, most of what hits your store is not shopping for you; it is feeding a model.

EngineCrawls per referral (Cloudflare Radar, early 2026)
Google~5:1
Perplexity~111:1
OpenAI (ChatGPT)~1,276:1
Anthropic (Claude)~23,951:1

Crawl share is not referral share

This is where a lot of AI-traffic writing quietly cheats, so it is worth stating plainly: the engine that crawls you the most is not necessarily the one that sends you the most humans.

Our first-party data measures crawl: ChatGPT’s crawler was 74 to 79% of recognized AI fetches. Referral share is a different metric, measured differently by different panels: Similarweb via Digital Applied (June 2026) puts ChatGPT at 52.7% of AI referrals with Gemini surging to 27.3% and Claude at 8.9%, while StatCounter (March 2026) places ChatGPT much higher at 78%. Treat the exact percentage as a range, not a fact. The robust, convergent signal is the direction: ChatGPT still number one but declining in relative share, Gemini number two and climbing fast.

The practical consequence: optimizing only for ChatGPT is a bet on a shrinking majority. Measuring your presence across engines, not just one, is now the safer posture.

Zero-click, and the France-versus-US tell

The last piece came from Search Console, and it is the cleanest natural experiment we had. For the same content, over 90 days, at nearly identical Google position (10.5 in the US, 10.6 in France), the US click-through rate was 0.08% (9 clicks on 11,936 impressions) while France was 4.63% (59 clicks on 1,274 impressions).

Nearly sixty times the CTR at the same rank is not a content difference. It is AI Overviews, live in the US and absent from France for the whole measurement window, blocked by neighboring-rights law. That window is now closed: Google rolled out AI Overviews and AI Mode in France on July 22, 2026 (Google France, July 22, 2026). The natural experiment is no longer reproducible, and reading it now becomes a forecast: if the mechanism is the one described here, French CTR at constant position should converge toward the US regime over the coming months. The rollout is progressive and limited to queries Google deems precise or complex enough, so expect gradual convergence rather than a cliff. In the US, Google increasingly answers inside the SERP, so the impression logs and the click never happens. We are, in effect, watching the zero-click effect of AI Overviews inside our own data.

What to actually do with this

Three takeaways survive the small sample, because each is corroborated by independent 2026 data:

  1. Measure the input, not just the output. The crawl (server or edge logs) is where you see the AI layer at all. If you only watch GA4, the AI layer is nearly invisible to you. See measuring AI traffic beyond GA4 for the setup.
  2. Recover the dark traffic. A large share of AI-referred humans hides in Direct. A custom channel grouping plus server-side signals brings some of it back into view.
  3. Get read, not just crawled. Being fetched 500 times a day is worthless if the engine cannot extract your price, your proof and your answer from the raw HTML. That is the difference between being crawled and being recommended, and it is exactly what a free GEO audit checks on a Shopify store.