Marketers increasingly talk about “AI visibility” as though ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode were five windows onto roughly the same ranking system. A large new dataset from Wellows suggests that assumption is badly misleading at the citation level. Across 22.7 million citations collected from January through June 2026, the five engines overwhelmingly drew on different sources even when they were answering the same questions.
The headline number from the Wellows AI Citation Overlap Study, updated August 6, is 79.6%. That is the share of website-and-question source pairings that appeared on only one of the five engines in the study’s all-engine comparison. Only 0.31% appeared across all five. In other words, being a source for one AI answer engine is far from equivalent to becoming a universal source across AI search.
22.7 million citations reveal five different source ecosystems
Wellows analyzed 22.7 million citations from 1.15 million questions and 441,946 websites across 27 markets, with 84% of the questions coming from the United States. For its strict five-engine comparison, the company narrowed the analysis to 531,889 questions on which ChatGPT, Gemini, Perplexity, Google AI Overviews and Google AI Mode all returned at least one cited website. That subset contained 13 million citations and 280,245 websites, allowing the researchers to compare the engines question by question rather than penalizing a platform for queries it did not answer with sources.
The resulting fragmentation is striking. Of 8,729,964 website-and-question pairings in the all-five dataset, 6,949,766 were cited by exactly one engine. Another 14.5% reached two engines, 4.4% reached three and 1.2% reached four. Just 27,097 pairings, or 0.31%, appeared across all five engines.
The pattern was also persistent rather than the product of one unusual month. Wellows reports that the single-engine share stayed between 77.7% and 81.0% throughout the six-month period. June showed somewhat more overlap than the pooled average, but the broader result remained intact: most sources selected for a given question were exclusive to one engine.
ChatGPT and Perplexity barely share the same source pool
The contrast becomes even clearer when the study isolates individual engines. Wellows found that 76.3% of sources cited by ChatGPT were not cited by any of the other four engines for the same question. Perplexity followed at 69.9%, Gemini at 63.7%, while Google AI Overviews and AI Mode were less unique at 50.4% and 49.7% respectively.
Pair ChatGPT directly with Perplexity and the separation is larger still. Perplexity did not use 89.1% of the websites ChatGPT cited for matching questions. The reverse comparison was almost symmetrical: ChatGPT did not use 90.1% of the websites cited by Perplexity. Wellows found similarly large gaps between ChatGPT and the other engines, including Gemini and Google’s two AI search surfaces.
That does not necessarily mean one engine is more erratic or less reliable. Wellows stress-tested parts of its analysis and concluded that much of the observed difference is better understood as a candidate-pool issue: the systems appear to draw from substantially different slices of the web. The study specifically revised an earlier interpretation that engines owned by the same company necessarily agree because they choose sources in similar ways. After adjusting for candidate-pool overlap, that ownership-based advantage largely disappeared.
A single AI visibility score can hide the information marketers need
This creates a measurement problem for SEO and GEO teams. If a dashboard compresses five engines into one AI visibility score, a rising aggregate can look reassuring while concealing which engine actually improved. A brand might be gaining citations in Google AI Overviews while remaining absent from ChatGPT, or improving in Perplexity while losing ground in Gemini.
Wellows therefore argues that any headline metric should open into engine-level reporting. Its broader AI visibility methodology tracks citations across the five platforms separately, reflecting the idea that the source URLs retrieved by each engine are an actionable layer that should not be treated as interchangeable.
There is an important caveat. The study measured overlap between source lists; it did not test whether a brand’s visibility score on one engine statistically predicts its visibility score on another. Two engines could cite different pages yet still produce correlated brand-level visibility over time. So the strongest defensible conclusion is not that a ranking or visibility result in one engine contains literally zero information about another, but that source-level success on one platform cannot safely be assumed to carry over to the other four.
Brands travel better than individual pages
The study also found a meaningful difference between page overlap and brand overlap. The engines landed on the same exact page only 6.8% of the time, while the same company appeared at a 30.3% rate. At first glance that might suggest brand authority transfers more cleanly between engines than page-level authority.
Wellows cautions against making that leap. Its permutation testing found that the raw 4.5-times gap between brand and page co-occurrence is partly explained by the structure of the comparison itself: an answer may name a relatively small set of brands while having thousands of possible pages available to cite. The company says brand overlap remains a useful practical observation, but not proof that brand-level signals inherently transfer better between engines.
For publishers, however, the distinction still matters operationally. Optimizing one “hero” article and expecting it to become the canonical AI source everywhere looks increasingly fragile. A broader strategy—building recognizable expertise around a topic, publishing multiple useful assets and earning relevant third-party references—creates more opportunities to enter the different candidate pools used by different systems.
GEO is becoming an engine-by-engine discipline
The practical consequence is that AI search optimization increasingly needs to be diagnosed by engine. If ChatGPT cites a site and Perplexity does not, rewriting the successful page may not address the real problem. The missing visibility could stem from differences in retrieval sources, publisher ecosystems, freshness, indexing, query expansion or other engine-specific mechanisms.
That also changes competitive reporting. “We rank in AI” is becoming too broad to be useful without specifying where. Teams need to know which prompts matter commercially, which engines their audience actually uses, which domains those engines cite for those prompts and whether improvements persist across repeated measurements rather than appearing in a single screenshot.
The Wellows dataset does not establish that the five systems will remain this fragmented forever. The company observed some increase in agreement from January to June, and AI retrieval systems are changing rapidly. But with nearly four out of five cited sources currently appearing on only one engine in the study, the burden of proof has shifted. Marketers should not assume cross-engine visibility; they should measure it.
For SEO teams accustomed to one dominant search index, that is the larger lesson. AI search is not yet behaving like a single rankings market with five interfaces. At the citation layer, it looks much more like several overlapping discovery systems with sharply different source pools. Winning one of them is useful. Treating that win as evidence that the other four have been won is not.