Marketers often talk about “ranking in AI search” as if ChatGPT, Perplexity, Gemini and Google AI Overviews were four interfaces sitting on top of roughly the same source hierarchy. A large cross-engine citation study suggests almost the opposite. The systems can answer the same question while drawing on remarkably different parts of the web.
In The Answer Index, published by Hendricks in September 2026, researchers analyzed 16,069 citations generated across four AI answer engines using a fixed corpus of 480 questions spanning eight industries. The strongest cross-engine pairing—Perplexity and Google AI Overviews—showed mean per-question domain overlap of just 0.181. In other words, even the two engines with the most similar citation sets shared only about 18.1% overlap by this measure.
Every pairing involving ChatGPT was dramatically lower. ChatGPT and Perplexity recorded mean overlap of 0.020, ChatGPT and Gemini 0.014, and ChatGPT and Google AI Overviews just 0.007. The finding challenges a convenient assumption behind many AI-visibility strategies: there may be no single universal source ranking that brands can optimize once and expect to carry across every answer engine.
The same question does not produce the same source set
The Answer Index was designed to make cross-engine comparisons more meaningful by holding the questions constant. The study used the same ten question types in each of eight industries, covering areas such as provider discovery, comparisons, cost, definitions and explanations. Each engine therefore received the same underlying research corpus rather than being compared through unrelated prompts.
That makes the low overlap notable. If source selection were mostly driven by a common web-wide authority hierarchy, the same question would be expected to surface many of the same domains across engines. Instead, the measured citation sets diverged substantially.
Perplexity and Google AI Overviews were the closest pair at 0.181 mean overlap. Gemini and Google AI Overviews followed at 0.125, while Perplexity and Gemini reached 0.106. Those figures are higher than the ChatGPT pairings but still indicate that most cited domains were not shared on the average question.
ChatGPT sits especially far from the other engines
ChatGPT’s results require an additional qualification because it cited sources far less frequently than the other systems in this experiment. Hendricks recorded at least one citation in only 72 of ChatGPT’s 480 answers. Perplexity cited sources in all 480, Gemini in 476 and Google AI Overviews in 430 of the 462 cases where an AI Overview rendered.
That lower citation frequency naturally affects the opportunity for overlap. But it is also part of the strategic difference. ChatGPT was not simply choosing a slightly different collection of domains; for many question types it was not surfacing citations at all. When it did cite, its source set showed very little agreement with the other engines.
The study found that ChatGPT’s citations were heavily concentrated in provider-discovery and cost questions. That means a brand’s apparent ChatGPT citation performance can depend on both whether the system retrieves sources for the prompt and whether the brand’s domain is selected once retrieval occurs. Those are two separate gates that may behave differently elsewhere.
Four answer engines should not be treated as one SERP
Traditional SEO trained marketers to think in terms of one dominant search index and one principal ranking environment. Rankings vary by location, device and personalization, but a page that performs strongly in Google generally exists within the same underlying search ecosystem as another Google query.
AI answer engines introduce a more fragmented structure. They can use different retrieval systems, source-selection logic, freshness mechanisms, model behavior and citation policies. Google AI Overviews has access to Google’s search infrastructure. Perplexity is built around source-backed answer generation. ChatGPT can answer some prompts from model knowledge and retrieve the web for others. Gemini has its own product and retrieval behavior.
The result is better understood as multiple source pipelines than a single ranking with four presentation layers. A domain that is highly visible to one engine may be nearly absent from another even when users ask equivalent questions.
Self-consistency is much higher than cross-engine agreement
One of the strongest pieces of evidence in the study is the difference between cross-engine overlap and same-engine repeatability. Hendricks reran a subset of 160 questions for Perplexity and Gemini to test how much each system agreed with itself across repeated runs.
Perplexity’s mean self-overlap reached 0.735, while Gemini’s reached 0.456. Those values are far above the 0.106 mean overlap measured between Perplexity and Gemini on the same questions. The study therefore argues that the engines are not merely producing random citation sets that happen to differ. Each appears to have a comparatively stable source-selection fingerprint of its own.
That distinction matters. If citation lists were simply unstable from one run to another, cross-engine differences would be less strategically meaningful. Higher self-consistency suggests that at least part of the divergence reflects systematic differences in how the engines retrieve or select sources.
There is no obvious universal “AI authority” score
The findings complicate attempts to build a single metric for AI-search authority. A domain can be frequently cited by Perplexity yet rarely appear in ChatGPT. Another can perform strongly in Google AI Overviews while having limited visibility in Gemini. Combining all four into one aggregate score can conceal those differences.
That does not make aggregate visibility useless. A cross-engine measure can still summarize broad exposure. But it should not be mistaken for a ranking signal shared by every platform. The underlying source pools can be sufficiently different that the aggregate is better thought of as a portfolio of engine-specific outcomes.
For brands, this resembles diversification more than conventional rank tracking. Success in one engine provides evidence that a source is useful in that ecosystem, but the Hendricks data gives little reason to assume that the same visibility will automatically transfer to the others.
Optimization may need to follow the source ecosystems
If the engines draw from different source pipelines, AI visibility strategy needs to examine more than a brand’s own website. Teams should identify which domains repeatedly appear for commercially important prompts in each engine. Those may include provider websites, publishers, directories, communities, review platforms or other intermediaries.
The relevant ecosystem can change by industry as well. The Answer Index found that provider-owned websites captured a large share of citation mass in some sectors, including healthcare and B2B SaaS, while education showed much greater reliance on other source types. Engine behavior and market structure therefore interact.
A useful audit might consequently produce four source maps rather than one. Which domains does ChatGPT cite when it retrieves? Which ones recur in Perplexity? Which sources dominate Gemini? Which publishers and provider sites appear in Google AI Overviews? The overlaps are worth noting, but the differences may reveal more actionable opportunities.
One piece of content cannot be assumed to travel everywhere
This fragmentation also changes expectations around content optimization. A page that becomes a frequent citation in one engine has clearly achieved something valuable, but marketers should resist treating that result as proof of universal AI-search visibility. The source-selection mechanisms may reward different signals or discover the page through different pathways.
That does not necessarily mean creating four versions of every article. It means measuring distribution separately. A strong primary source, useful original data and clear factual structure can have value across systems, but teams still need to verify where the content actually surfaces rather than extrapolating from one platform.
Third-party visibility can be equally engine-specific. If an authoritative publisher is repeatedly cited in one answer engine but absent from another, earning coverage there may disproportionately improve visibility in the first ecosystem. Digital PR, content partnerships and source outreach could therefore become increasingly informed by engine-level citation data.
AI search looks less like one ranking and more like four markets
The Answer Index is not a universal census of AI search. Its eight industries are a controlled sample, and citation behavior can change as models, retrieval systems and products evolve. The study itself presents the framework as something companies should replicate with questions specific to their own markets.
But the scale of the divergence is difficult to ignore. Across 16,069 citations, the most similar cross-engine pair averaged only 18.1% domain overlap per question. Every pairing involving ChatGPT remained at 2% or below. Meanwhile, repeated runs of individual engines showed much stronger internal agreement.
For marketers, that suggests a more useful mental model for AI search. There is not necessarily one hidden ranking waiting to be decoded. There are multiple systems with their own retrieval behavior, citation frequency and source preferences. Winning visibility in one is valuable, but it does not mean the other three are seeing the web the same way.
The practical implication is simple: measure AI engines separately before combining them. A single blended visibility score can be a dashboard convenience, but the strategic work happens underneath it—in the four different source pipelines that decide which parts of the web each answer engine chooses to trust.