Your AI Visibility Dashboard May Be Measuring the Wrong Reality: Chatbot Interfaces and APIs Barely Share the Same Sources

Your AI Visibility Dashboard May Be Measuring the Wrong Reality: Chatbot Interfaces and APIs Barely Share the Same Sources
Sponsored

AI visibility dashboards are increasingly being sold as the generative-search equivalent of rank trackers: enter a set of prompts, monitor which brands and domains appear, and use the resulting citation data to judge performance. A new academic preprint suggests that this model may be measuring only one version of reality.

Researchers auditing commercial product recommendations across ChatGPT, Google Gemini and Google AI Overviews found striking differences not only between competing AI products, but between consumer-facing chatbot interfaces and the APIs that developers often use to measure them at scale. The study, submitted to arXiv on September 16, 2026, analyzed 1,536 responses generated from a dataset of 2,528 real commercial-advice queries. It has not yet been peer-reviewed, so its findings should be treated as preprint evidence rather than settled conclusions.

The practical implication for generative engine optimization is significant. If an AI monitoring platform collects responses through an API while actual customers use a consumer chatbot interface, the two environments may expose substantially different sets of sources. A clean dashboard can therefore be internally consistent while still failing to represent what users actually see.

ChatGPT and Gemini rarely displayed the same domains

The researchers created ConsumerQ, a dataset of 2,528 real commercial-advice queries, and evaluated product-related responses from ChatGPT through both its chatbot and API, Gemini through its chatbot and API, and Google Search AI Overviews. Their audit was designed to examine how AI systems recommend products and what source information they expose.

For the same query, the consumer ChatGPT and Gemini interfaces shared only 5.4% of displayed domains on average. More strikingly, 76.7% of comparisons had no displayed domain in common at all. In other words, two users asking comparable commercial questions across the two major chatbot ecosystems could encounter almost entirely different source landscapes.

This reinforces a broader pattern already visible in generative search research. NetContentSEO recently covered a separate 11,500-query study in which Google Search and AI Overviews shared only 18% of their sources. Together, the studies point toward a visibility environment fragmented not just by ranking position, but by product, interface, retrieval system and response generation.

The API may not represent the chatbot your customers use

The most consequential finding for AI visibility vendors concerns the gap between interfaces and APIs. According to the preprint, mean domain overlap between ChatGPT’s API and its corresponding consumer interface was just 12%. For Gemini, the mean overlap was 14.8%. The authors also report differences in the types and layers of source information exposed by the two environments.

Those numbers do not mean APIs are useless for measurement. APIs remain essential for scalable experiments, repeatable testing and controlled comparisons. What the results challenge is the assumption that an API observation can automatically be treated as a faithful proxy for a consumer-facing answer.

This distinction matters because GEO platforms face an engineering tradeoff. Consumer interfaces can be difficult to monitor reliably at large scale, while APIs are structured, automatable and easier to query repeatedly. But if the measurement layer changes the observed source ecosystem, scale alone does not solve the validity problem. A dashboard may precisely measure API visibility while labeling it more broadly as ChatGPT or Gemini visibility.

One AI response is not a stable ranking

The study raises another problem for rank-tracker-style thinking: repeated prompts do not necessarily return stable recommendations. The researchers report that recommended products often changed across repeated requests. This means a single answer should not be interpreted in the same way as a deterministic search ranking captured at one moment.

For AI visibility measurement, the unit of analysis may therefore need to become a distribution rather than an individual response. Instead of asking whether a brand appeared for a prompt, a more useful system would ask how frequently it appeared across repeated runs, which sources accompanied it, how those sources varied by interface, and whether the result persisted over time.

That interpretation fits earlier evidence of citation instability. In another analysis covered by NetContentSEO, roughly half of the observed AI citation sources in one GEO experiment disappeared within 30 days. Citation visibility increasingly looks less like a fixed SERP position and more like a changing probability distribution.

ChatGPT showed far more first-person product preferences

The audit also found a large behavioral difference in how the systems framed product recommendations. Among responses that recommended products, ChatGPT expressed a first-person product preference in 79% of cases, compared with 7% for Gemini and 2% for Google AI Overviews.

This is an important finding, but it should not be overextended. The study focuses on commercial product-advice queries and does not establish that ChatGPT behaves this way across all informational or navigational searches. Nor does the percentage by itself demonstrate commercial bias toward any particular brand. It shows that, within the audited recommendation context, the systems used markedly different styles of recommendation.

For marketers, however, presentation style can influence what “visibility” means. Being listed among several alternatives is not equivalent to being incorporated into a direct recommendation, just as appearing in a source panel is not necessarily equivalent to shaping the generated answer. NetContentSEO recently examined this distinction through research proposing an “absorption” metric for measuring whether cited content actually contributes to an AI response.

AI visibility needs to specify what is being observed

The emerging measurement problem is methodological. A report saying that a domain has “20% ChatGPT visibility” is difficult to interpret unless it specifies the environment, model or product surface, prompt set, geography where relevant, number of repetitions, observation period and definition of visibility. The new study adds another field to that list: whether the observation came from the consumer interface or an API.

This does not make AI visibility measurement impossible. It makes transparent methodology more important. API-based monitoring can still provide useful trend data, particularly when the same controlled method is applied consistently over time. Consumer-interface testing can provide a closer view of the experience users encounter. Combining both can reveal whether movements in a monitoring system correspond to movements in the public product.

For serious GEO programs, a stronger design would therefore treat API and interface observations as separate measurement layers rather than silently merging them. It would run repeated prompts instead of relying on one response, track source-domain distributions rather than only binary citations, and distinguish between being displayed as a source, being mentioned in the answer and materially influencing the recommendation.

The dashboard is a model of reality, not reality itself

The preprint’s strongest lesson is not that existing AI visibility tools are wrong. It is that their outputs are conditional on how they observe the system. The researchers explicitly conclude that neither isolated responses nor API observations should automatically be assumed to represent the commercial advice consumers encounter.

That qualification should become standard in GEO reporting. Traditional SEO spent decades learning that rankings vary by location, device, personalization, query interpretation and time. AI search adds further dimensions: stochastic generation, retrieval variation, interface differences, model changes and source-layer differences between APIs and consumer products.

A dashboard can still be extremely useful, but its numbers need a clearly defined denominator. If ChatGPT’s API and consumer interface share only 12% of source domains on average in this audit, measuring one and naming it as the other risks false precision. The next generation of AI visibility tools will need to show not only where a brand appears, but which version of the AI ecosystem was actually measured—and how closely that environment resembles the one real customers are using.

0%