Your Page Can Influence an AI Answer Without Receiving a Citation: Retrieval Visibility Is Larger Than Visible Citation Visibility

Your Page Can Influence an AI Answer Without Receiving a Citation: Retrieval Visibility Is Larger Than Visible Citation Visibility
Sponsored

A webpage can contribute evidence to an AI answer without ever appearing in the citations the user sees. That is the most important measurement implication of a new academic study of conversational web search across ChatGPT, Claude, Grok and DeepSeek. The researchers found that agents frequently retrieve far more URLs than they ultimately cite—and, more importantly, that some claims in the final response can be supported by those retrieved-but-uncited pages.

The study, “Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses,” was submitted to arXiv on September 16 by researchers affiliated with institutions including the Max Planck Institute for Software Systems, Saarland University, Ruhr University Bochum, Seoul National University, École Polytechnique and Microsoft. It is a preprint rather than a peer-reviewed publication, so its results should be treated as substantial new evidence about the measured systems, not as a universal law of AI search.

Its scale is unusual. The researchers analyzed 171,264 donated real-world conversations from 613 users across the four platforms, with fine-grained traces containing prompts, web queries, search results, cited URLs and generated responses. They then complemented those observational data with controlled API experiments in which frontier models received the same 1,000 prompts. That combination lets the paper examine something citation dashboards normally cannot see: the path between the search query, the pool of returned URLs and the sources ultimately exposed to the user.

Claude searched 825 times; GPT searched only 140

The first major divergence appears before citations even enter the picture. In the controlled experiment, Claude Sonnet 4.6 invoked web search for 825 of the 1,000 prompts, while GPT-5.3-chat searched for only 140. Grok-4.3 searched on 766 prompts and DeepSeek-v4-flash on 584. The researchers conclude that conversational agents make substantially different decisions about when external information is necessary.

More searching did not automatically produce better answers. The paper reports that web search generally improved response quality when GPT, Grok and DeepSeek chose to invoke it, while Claude’s much more frequent search behavior produced mixed effects across the measured quality dimensions. The point is not that searching less is inherently better; it is that retrieval frequency and answer quality are separate variables.

For publishers, this means AI visibility begins even earlier than source selection. Before a page can influence an answer, the platform must decide to search at all. Different agents can receive the same prompt and create radically different opportunities for external websites simply because one opens a retrieval process while another relies on parametric knowledge.

Only a fraction of retrieved URLs become visible citations

The study then measures what happens to URLs after search. In the real-world conversation traces, ChatGPT returned 1,196,238 search-result URLs but cited 159,580 of those retrieved URLs, corresponding to a 13.3% citation rate under the paper’s definition. Claude’s rate was 19.9%, Grok’s just 1.8% and DeepSeek’s 34.1%. In other words, the visible citation layer represented only a subset of the web pages that entered the observable retrieval pipeline.

The controlled experiments produced different percentages—10.6% for GPT-5.3-chat, 39.7% for Claude Sonnet 4.6 and 25.5% for Grok-4.3—reinforcing another warning for AI visibility measurement: interface behavior and API experiments are not interchangeable. The exact rates depend on the platform, model, harness and experimental setting.

This provides direct empirical support for a distinction NetContentSEO has repeatedly highlighted. In our analysis of 1.6 million AI citations, we noted that a citation report observes the URL ultimately exposed as a source, not every intermediate resource the system may have discovered or retrieved. The new preprint goes further because it observes search traces directly and shows just how large that hidden retrieval set can be.

Some uncited pages still support claims in the answer

The more consequential finding comes from the paper’s grounding analysis. The researchers decomposed responses into atomic claims and tested whether each claim was entailed by content scraped from cited and retrieved URLs. Their automated entailment judgments were validated against human annotations, with 93% agreement in the sampled validation exercise.

Across ChatGPT, Claude and Grok, 14% to 53% of claims across the real-world and controlled settings were supported by URLs that had been returned during search but were never cited in the final response. These pages were not merely sitting unused somewhere in the candidate pool: their content could support claims that appeared in the answer even though the user was not shown them as sources.

The authors frame this primarily as an attribution problem. If an agent uses evidence from a page without citing that page, the original source may not receive appropriate credit. For publishers and GEO teams, however, it is also a measurement problem. Citation visibility can underestimate informational influence because the observable link list is smaller than the set of retrieved pages that may contribute evidence.

A citation dashboard measures only one layer of AI visibility

Most commercial AI visibility systems are necessarily built around what can be observed from the outside: brand mentions, cited URLs, source domains, answer text and sometimes referral traffic. Those are valuable metrics, but they largely describe the final response layer. The new research suggests that an additional retrieval layer exists upstream and can matter even when it leaves no visible attribution trace.

This creates at least three distinct states for a publisher. A page can fail to enter retrieval at all. It can be retrieved and potentially contribute evidence without receiving a citation. Or it can be retrieved, selected and visibly cited. Treating only the third state as “AI visibility” collapses very different outcomes into the same apparent zero whenever a URL is absent from the source list.

That distinction complements another recent NetContentSEO analysis, which examined the opposite problem: a page can be cited without strongly shaping the answer. Put the two findings together and citation counts become an increasingly imperfect proxy for influence. An uncited page may help support the answer, while a cited page may contribute relatively little to its substance.

Retrieval visibility and citation visibility should be measured separately

A more useful GEO funnel therefore begins with retrieval rather than citations. First, can the system discover and fetch the page? Second, does the page enter the evidence pool for relevant prompts? Third, does its content support or shape the generated response? Fourth, does the system visibly attribute that contribution through a citation? Finally, does the citation or brand mention generate user action?

Most publishers can directly observe only parts of this sequence. Server logs may reveal some crawler or fetch activity, but they do not necessarily expose which prompt triggered it or whether retrieved content entered a particular answer. Citation monitoring reveals visible attribution but misses uncited retrieval. Referral analytics captures clicks but says little about answers that influenced users without generating a visit.

This is why a simple diagnostic for retrieval can still be useful. NetContentSEO recently covered a workflow for testing whether an AI assistant can retrieve exact content before chasing citations. Such tests cannot reveal proprietary indexes or prove that a page will influence future answers, but they isolate an earlier stage of the funnel that citation monitoring skips.

The platforms also retrieve from very different parts of the web

The study finds another complication: the search engines behind conversational agents show different domain preferences. In the real-world traces, the top ten domains represented 21.3% of ChatGPT’s returned results and 32.3% of Grok’s. Reddit and YouTube were prominent among results returned for ChatGPT and Grok but absent from Claude’s observed search results, while DeepSeek showed its own distinctive source preferences.

These results should not be converted into a permanent list of domains that a publisher “must” target. The study covers particular platforms, users and periods, and the authors explicitly warn that proprietary retrieval pipelines and ranking algorithms remain hidden. But it does show that the retrieval market itself differs by platform before citation selection begins.

For GEO measurement, that means two systems can disagree about a publisher for multiple reasons. One may never retrieve the page. Another may retrieve it but not cite it. A third may cite it without relying heavily on its content. A single share-of-citations number cannot distinguish among those mechanisms.

The 171,264 conversations do not represent every AI user

The scale of the dataset is impressive, but its limitations matter. The 171,264 conversations came from 613 consenting users rather than a representative sample of every ChatGPT, Claude, Grok or DeepSeek user. The paper primarily examines English-language text interactions and does not expose proprietary ranking algorithms, hidden reasoning traces or the full internal retrieval architecture of the platforms.

Several analyses also rely on automated judge models, although the researchers performed human validation and report high agreement for the grounding evaluation. The controlled experiments use specific 2026 model versions and API configurations, meaning future product changes could alter search and citation behavior. The authors themselves caution that some observed differences may reflect platform harnesses rather than the underlying language models alone.

Those limitations do not erase the central measurement result. Within the observed traces, retrieval and visible attribution were demonstrably different layers. Large numbers of URLs entered search results without becoming citations, and a meaningful fraction of answer claims could be grounded in those uncited retrieved pages.

GEO needs a metric for influence before attribution

The emerging picture is more complicated than “rank in AI answers” or “win more citations.” AI search is a pipeline of decisions: whether to search, what query to formulate, which pages to retrieve, which evidence to use, which sources to cite and which links a user ultimately follows. A publisher can succeed or fail at each stage independently.

Visible citations remain valuable because they provide attribution, brand exposure and a possible path to referral traffic. But they should not be mistaken for a complete inventory of the sources influencing an answer. The new study provides unusually strong evidence that the hidden retrieval layer is larger than the visible citation layer—and that some of what disappears from the source list may still survive in the answer itself.

For publishers, that changes the question from “Did the AI cite us?” to a harder sequence: can the system retrieve us, does it use our evidence, does it credit us, and does that attribution reach the user? Citation tracking measures one of those outcomes. The next generation of AI visibility analytics will need to understand the rest.

0%