Something Inc.Schedule a free consultation
RESEARCH

The AI discovery funnel: where citations leak before they become traffic

Correct ai traffic attribution requires tracking three separate leaks — crawl, citation, and click — and most GEO programs still measure only the middle one.

3 DATASETSGEOJUL 2026

Most GEO reporting tracks one number: citations. A dashboard shows a brand got mentioned in ChatGPT or AI Mode some number of times this month, the line goes up, and the program gets called a win. That number is real, but it sits in the middle of a funnel that starts with a crawler request and ends, or does not end, with a human on a page. Correct ai traffic attribution means tracking all three stages — crawl, citation, and click — because each one leaks independently, and the leaks do not add up the way most reporting assumes.

This piece cross-references three independently published July 2026 datasets to build that funnel model: how much of a crawled site ever sends traffic back, how much of what gets crawled actually earns a citation, and how much of the traffic tied to a citation actually lands on the page that earned it. The answer at every stage is: less than most programs assume, and the gap changes what a GEO report should be measuring.

The reason this matters beyond methodology is financial. Budget gets allocated based on what gets measured, and citation counts are the easiest of the three stages to put in a slide. Crawl access lives in server logs most marketing teams never open. Traffic landing pages live in an analytics view most GEO tools do not connect to citation data at all. So the metric that is easiest to report becomes the metric that drives the budget conversation, even though it is the one stage in the funnel where a brand has the least direct control over the outcome once the content itself is built. Fixing that ordering, measuring crawl and traffic with the same rigor already applied to citations, is the practical output of this analysis.

TL;DR · 60 SECONDSThe AI discovery funnel leaks at three separate stages, and most GEO measurement only sees the middle one. Stage 1, crawl: ClaudeBot requested 20,583 pages for every one visit it referred back in Q1 2026, versus Googlebot's roughly 5:1, per TechnologyChecker.io's analysis of Cloudflare Radar data (David Thomson). Stage 2, citation: external, third-party sources account for 84% to 93% of an AI engine's citation weight depending on vertical and platform, and ChatGPT and Google AI Mode weight evidence types so differently — video is ~1% of ChatGPT's external evidence versus 23% of AI Mode's — that optimizing for one engine can actively hurt visibility on the other, per Aleyda Solis's analysis of 15 SaaS brands. Stage 3, traffic: even after a citation lands, up to 67.3% of the AI-driven traffic in a vertical goes to pages that were never cited at all. Citation counts, on their own, cannot explain most of what a GEO program actually delivers or fails to deliver.
METHODOLOGYThis analysis cross-references independently published July 2026 datasets to construct a funnel model of how AI search traffic actually moves from crawl to citation to click. Something Inc. did not run new primary data collection for this piece — every figure below is sourced and attributed to its original publisher: TechnologyChecker.io's robots.txt and Cloudflare Radar analysis (David Thomson, published April 3, 2026, updated July 2, 2026) for the crawl stage, and Aleyda Solis's Semrush-and-Similarweb-based study of 15 SaaS brands (published July 16, 2026) for the citation and traffic stages. Each source retains its own methodology and time window. We did not re-run or independently verify their underlying data collection; we did verify every figure cited here against the original publication before using it.

The three-stage funnel nobody reports on

GEO measurement, as most agencies and in-house teams run it today, starts at the citation stage and stops there. A tool tracks how often a brand gets mentioned across a set of prompts, reports a mention rate or a share-of-voice number, and calls that AI visibility. It is a real signal. It is also the middle third of a longer process, and treating it as the whole process is why so much GEO reporting fails to reconcile with what pipeline actually shows. A citation is not the end of the funnel. It is one checkpoint inside it.

Before a page can be cited, an AI crawler has to request it, and most of those requests never turn into anything the site can measure — that is Stage 1, the crawl leak, and it happens entirely outside the page itself, dictated by robots.txt rules, server responses, and whatever a crawler decides is worth fetching. After a page is crawled and indexed by whatever retrieval system sits behind an engine, it still has to be selected as evidence for a specific answer, weighted against everything else the engine considered, and named — that is Stage 2, the citation leak, and it is the only stage most GEO tools currently instrument. And after a citation exists, the click that follows it does not reliably land on the cited page — that is Stage 3, the traffic leak, and it is the one that breaks pipeline attribution models built on citation counts alone.

HOW A QUESTION BECOMES A CITATION
CrawlPages requested by AI crawlers vs. pages that ever send a referral back
CitationOf what's crawled and indexed, the share actually named as a source in an answer
TrafficOf what's cited, the share of resulting clicks that land on the cited page itself
THE GAP THIS EXPLAINSA brand can improve its citation rate quarter over quarter and still see flat or falling pipeline, because the crawl stage determines what is even eligible to be cited, and the traffic stage determines whether a citation converts into a session on the page that earned it. Both are invisible to a citation-only dashboard.

Stage 1: the crawl leak

TechnologyChecker.io's David Thomson analyzed 4,047 robots.txt files against Cloudflare Radar's crawl and referral data, which spans infrastructure in 330 cities across 125-plus countries, to build what he calls a crawl-to-refer ratio: how many pages a bot requests for every visit it sends back to the site. The gap between AI crawlers and a traditional search crawler is not incremental. It is an order-of-magnitude difference in how much a site pays in server load and content exposure for how little it gets back in return.

In Q1 2026, ClaudeBot requested 20,583 pages for every one referral it produced. GPTBot's ratio was 1,255:1. Googlebot, by comparison, sat at roughly 5:1 — a crawler that still operates on something close to the old bargain, where being crawled reliably correlates with being found. Thomson's July 2 update adds a year-over-year lens that matters as much as the ratios themselves: both AI crawlers improved by Q2 2026 close, ClaudeBot's ratio narrowing to roughly 5,143:1 and GPTBot's to about 870:1, but improved is a relative term at that scale. A crawler that visits a page 5,000 times for every session it sends is still an economically strange visitor, and sites are responding accordingly.

ClaudeBot (20,583 requests : 1 referral)100%
GPTBot (1,255 requests : 1 referral)61%
Googlebot (5 requests : 1 referral)4%

Crawl-to-refer ratio, Q1 2026 — pages requested per one referral sent (TechnologyChecker.io / Cloudflare Radar, n=4,047 robots.txt files)

That gap is why blocking is rising. The share of sites returning a 403 Forbidden response to AI crawlers climbed from 3.63% to 8.56% year over year, more than doubling in twelve months. Some of that is deliberate policy — publishers and SaaS vendors deciding an AI crawler's request volume is not worth its referral value and shutting the door. Some of it is accidental: a CDN rule, a bot-management default, or a WAF policy written for generic scraper traffic that silently catches ClaudeBot or PerplexityBot along with it. Either way, the effect on a GEO program is the same. A page that returns a 403 cannot be crawled, cannot be indexed, and cannot be cited, regardless of how well it is written or how strong its comparison content is. The citation-stage optimization most GEO programs run assumes the page is reachable in the first place, and for a rising share of the web, in 2026, that assumption no longer holds by default.

This is also the stage where crawl economics stop being an abstraction. Cloudflare's own pay-per-crawl experiment, does it pay, was a direct response to exactly this ratio problem: if a crawler visits 20,000 times for one referral, metering that access starts to look less like friction and more like the only rational response available to a publisher. Whatever a site decides about paywalling or throttling AI crawlers, the crawl-to-refer ratio is the number that decision should be based on, and it is a number almost no GEO reporting currently surfaces.

Stage 2: the citation leak

Assume a page clears Stage 1: it is crawlable, indexed, and eligible to be cited. The next leak is inside the engine's own selection logic, and Aleyda Solis's July 16, 2026 analysis of 15 SaaS brands across CRM/Sales, Project/Collaboration, and Accounting/Finance verticals, using Semrush citation data and Similarweb AI referral data, quantifies it with more precision than most GEO research to date.

The headline finding: third-party, external sources dominate citation weight everywhere, but not evenly. External sources account for 92.1% of ChatGPT's citation weight and 93.4% of AI Mode's in CRM and Sales; 91.9% and 92.2% respectively in Project and Collaboration; and 83.6% and 84.1% in Accounting and Finance, the vertical where a brand's own domain carries the most relative weight of the three. In every vertical and on every engine, a brand's own pages are a minority of the evidence an AI answer draws on. Owned-content optimization, on its own, cannot fix a citation gap that is structurally weighted toward what other people publish about a brand.

VERTICALEXTERNAL WEIGHT, CHATGPTEXTERNAL WEIGHT, AI MODE
CRM & Sales92.1%93.4%
Project & Collaboration91.9%92.2%
Accounting & Finance83.6%84.1%

The second finding is the one most GEO programs are not built for at all: ChatGPT and AI Mode do not just cite different pages, they weight entirely different categories of evidence. Video content makes up roughly 1% of ChatGPT's external evidence but 23% of AI Mode's, a 23x gap that reflects Google's access to YouTube's index in a way OpenAI simply does not have. Peer-vendor citations run the other direction: 25.4% of ChatGPT's external evidence versus 16.1% of AI Mode's. Technology-publication citations show the widest swing of all in Solis's data, 7.7% for ChatGPT against roughly 0.5% for AI Mode. A GEO program optimizing toward a single blended "AI visibility" score is, in effect, averaging together two engines with close to opposite editorial preferences, which is a large part of why cross-engine citation strategies so often underperform single-engine ones.

Page type shows the same divergence. Homepages capture 13.7% to 16.9% of ChatGPT's citation weight across the three verticals, a meaningful share of a brand's own citation footprint. On AI Mode, homepages capture 0.4% to 4.0%. A brand that leans on its homepage's authority to win ChatGPT citations is building on ground that barely exists for Google's engine, and a program measuring "citations" as one undifferentiated total will not surface that a homepage-heavy strategy is quietly failing on half its target engines.

The platforms don't just cite different pages — they reward almost opposite kinds of evidence. — Aleyda Solis, SaaS AI search optimization analysis, July 16, 2026

None of this contradicts what comparison content earns in citation share at the format level; it sits one layer beneath it, at the level of source type and engine, and it is the layer a citation-count dashboard cannot see because a citation count does not carry a vertical, an engine, or an evidence-type tag by default.

Stage 3: the traffic leak

The most consequential leak in the funnel, and the one that breaks the most pipeline models, sits after the citation. A citation is not a session. Solis's traffic data, drawn from Similarweb referral tracking against the same 15-brand sample, shows that a meaningful share of the traffic AI engines send to a domain never touches the page that was actually cited.

In the Project and Collaboration vertical, 67.3% of AI-driven traffic landed on pages that were never cited at all — the engine named one page in its answer, and the click went somewhere else on the same domain entirely. CRM and Sales showed 42.7% of traffic on uncited pages. Accounting and Finance, the vertical with the tightest citation-to-traffic link, still saw 22.4% land off the cited page. Even in the best case in this dataset, more than one in five AI-referred visits cannot be explained by the citation a program is tracking.

Looking at it from the other direction sharpens the point. Restricting to ChatGPT specifically, the share of traffic that does land on the exact page cited varies from 76.6% in Accounting and Finance, down to 64.5% in CRM and Sales, down to just 39.8% in Project and Collaboration — meaning in that vertical, a clear majority of ChatGPT-driven visits go somewhere other than the page ChatGPT named as its source.

VERTICALTRAFFIC ON UNCITED PAGESCHATGPT TRAFFIC ON THE CITED PAGE
Project & Collaboration67.3%39.8%
CRM & Sales42.7%64.5%
Accounting & Finance22.4%76.6%
20,583:1
ClaudeBot's crawl-to-refer ratio, Q1 2026 (TechnologyChecker.io)
84-93%
citation weight held by external, third-party sources across verticals (Aleyda Solis)
67.3%
of AI traffic in Collaboration landed on a page that was never cited (Aleyda Solis)

Two mechanisms plausibly explain most of this. First, AI answers routinely cite one page while summarizing information that also lives on adjacent pages — pricing, feature, and integration pages that echo the cited page's claims — so a reader who clicks through to explore ends up browsing the site rather than staying put on the single URL the engine named. Second, engines cite specific deep pages, like a comparison or pricing page, but users landing from an AI answer frequently navigate to a homepage or a more general category page first, especially when the AI-generated summary primed them with a broad question rather than the narrow one the cited page answers. Either way, a pipeline attribution model that credits deals only to sessions starting on the exact cited URL will systematically undercount AI's real contribution, sometimes by more than half.

For scale context on the traffic stage: ChatGPT already accounts for roughly 92% of trackable LLM referral traffic industry-wide, per Previsible's analysis. That concentration is worth knowing when reading Stage 3's numbers, since ChatGPT's citation-to-traffic behavior is doing most of the work behind any blended AI-traffic figure a site sees in its own analytics.

What this means for GEO measurement

Put the three stages together and the shape of the problem is clear. A citation-only GEO report answers one question — did the brand get named — and treats that as a proxy for two questions it cannot actually answer: was the brand reachable in the first place, and did the citation produce a session on the page that earned it. Both of those questions have real, sourced answers now, and both point toward measurement gaps wide enough to explain most of the disconnect between rising citation counts and flat pipeline.

01Crawl access is a prerequisite metric, not an afterthoughtTrack crawl-to-refer ratio and 403 rate per AI crawler alongside citation rate. A rising 403 rate on ClaudeBot or GPTBot explains a falling citation count before any content problem does.
02One blended citation score hides two opposite enginesChatGPT and AI Mode weight video, peer vendors, and homepages differently enough that a single "AI visibility" number averages together strategies that should not be averaged. Report citation share per engine, per evidence type.
03Pipeline attribution needs a domain-level window, not a URL matchCrediting AI-driven pipeline only to sessions that start on the exact cited page will undercount the real contribution by double digits in some verticals. Attribution windows should credit the domain for a defined period after an AI referral, not just the single landing URL.

None of this argues for abandoning citation tracking — it is still the only stage of the three that a brand can influence directly and quickly through content decisions. It argues for putting citation tracking in its proper place: the middle instrument in a three-instrument panel, sitting between a crawl-access check that determines whether content is even eligible and a traffic-attribution model that determines whether a citation actually pays off. A GEO report built on citations alone is reading one gauge on a dashboard that has three, and the two it is ignoring are the ones that decide whether the whole program shows up in revenue.

This is also where reporting and analytics work earns its keep separately from content and technical GEO work. The crawl-access audit that catches an accidental block, the per-engine citation breakdown that stops one blended score from hiding a real gap, and the domain-level attribution window that credits pipeline correctly are three distinct measurement builds, not one dashboard widget. For SaaS teams specifically — the tech and SaaS buyers in Solis's underlying dataset — the vertical-level swings above are large enough that a program benchmarked against the wrong vertical's numbers will misjudge its own performance in either direction.

The mechanics of doing this well, from crawl-access checks through to generative engine optimization execution and a 30-day rollout, are covered in our 2026 AI discovery readiness playbook. And for a live example of what closing the traffic-leak gap looks like at the reporting layer, see how Zenity tied rankings, citations, and pipeline attribution into one weekly measurement loop instead of reporting citations alone.

Apr 3, 2026, updated Jul 2, 2026
TechnologyChecker.io — David ThomsonCrawl-to-refer ratios and 403 rates for AI crawlers, drawn from 4,047 robots.txt files and Cloudflare Radar data.
Jul 16, 2026
Aleyda Solis — SaaS AI search optimizationCitation-weight and traffic-landing data across 15 SaaS brands in three verticals, from Semrush and Similarweb.
Cite this research● LIVE
Something Inc. (2026). The AI Discovery Funnel: Where Citations Leak Before They Become Traffic.
Synthesis of TechnologyChecker.io/Cloudflare Radar crawl data and Aleyda Solis's SaaS AI search citation and traffic data.
somethinginc.com/insights
THE CLOSEEvery stage of this funnel is measurable today with tools most enterprise teams already own. The reason most GEO programs don't measure all three is that citation tracking shipped first and got mistaken for the whole picture. It isn't. Build the crawl-access check and the traffic-attribution window before the next budget cycle, or keep explaining a citation chart that never quite matches the pipeline number next to it.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.