AI Shopping Facts Clash in 86% of Tested Queries

AI Shopping Facts Clash in 86% of Tested Queries

Product.ai reports repeatable factual conflicts in 86% of the shopping questions it could score across tested AI configurations. The finding raises an important question for commercial AI visibility: when a product appears in an answer, are its identity, specifications and price represented correctly?

What the 86% figure measures

The company study, published September 22, 2026, collected 8,794 answers to 220 questions across ChatGPT, Claude, Gemini and Perplexity, with two configurations per provider and five repetitions. It scored 217 complete question groups; 187 had a confirmed conflict. Its definition requires a checkable factual disagreement recurring in at least two runs, rather than an opinion difference.

Comparison questions had a 97% conflict rate. In a separate price check, 913 answers were verifiable: 85% matched the accepted seller prices. Among wrong prices, the median miss was $300. That is not the median error across all answers.

These are API tests collected September 1, not direct consumer-app tests. Claude served as scoring judge, with anonymisation and additional checks disclosed. This is research published by a commerce-verification company, not an independent academic trial.

Google challenges the consumer-app comparison

Business Insider’s Japanese edition, carrying Alex Bitter’s reporting, says Google questioned the use of the Gemini API rather than its consumer app. Google said the app uses its Shopping Graph and continuously updated product listings.

The response identifies an important boundary in what was tested. It does not independently establish the consumer app’s accuracy. Equally, the API findings do not measure every shopper’s experience. Evaluating either claim would require observations from the relevant product under clearly recorded conditions.

A conflict rate is not an answer-level error rate

There is a mathematical distinction between counting questions with a conflict and counting individual wrong answers. A question asked many times can qualify for the first measure even when most answers agree with a verified reference. Without an answer-level denominator and a claim-by-claim check, the question-level percentage cannot be converted into the probability that any single recommendation is wrong.

Agreement is also insufficient on its own. Two answers can repeat the same incorrect specification. Two different prices can both be legitimate if they describe different sellers, variants or dates. A useful investigation needs to preserve those conditions rather than treating every difference as the same failure.

For NetContentSEO, the implication is that visibility and factual fidelity should be reported separately. A brand mention records exposure. It does not tell a business whether the answer matched the right product or preserved the commercial conditions of an offer.

What commercial teams should record

Our proposed reporting approach begins with an exact product identifier and a dated reference. Keep the model or variant, seller, currency, package quantity and price conditions together. For specifications, record the measurement context: a battery-life claim under one test setting may not be comparable with another. These are editorial recommendations, not additional results from Product.ai’s experiment.

When an answer conflicts with that reference, preserve the actual claim and its linked evidence. Categorise whether the issue concerns product identity, specification, price or availability. Keep unresolved cases visible; an inaccessible source should not automatically become a confirmed factual error.

Consider a hypothetical listing for a laptop with two storage configurations. If an answer names the larger configuration but supplies the smaller model’s price, counting the brand mention as a success would conceal the mistake. The corrective task is to investigate where the identity and price became separated, rather than simply trying to obtain more mentions.

A testable response to inaccurate representation

A retailer or manufacturer could use such records to review its own public product information. Check whether variant names are consistent, current offers carry clear conditions and obsolete pages distinguish historical facts from current availability. Keep dated copies before changing content so later comparisons have an inspectable baseline.

Any claimed improvement should then be tested against the same defined task. A cleaner catalogue may be useful, but it does not guarantee that an external assistant will retrieve the updated page or reproduce it correctly. Recording successes, failures and missing evidence makes that limitation measurable.

Product.ai’s study offers a concrete signal to investigate, with a narrower scope than a universal verdict on AI shopping. For businesses, the useful response is to measure whether their products are described accurately alongside how often they appear.

0%