AI product recommendations are becoming a discovery layer between consumers and commerce, but a single answer is a weak measurement of how that layer behaves. A recent empirical audit found substantial variation in recommended products, recommendation language and displayed sources when commercial queries were repeated across consumer interfaces and APIs.
The researchers assembled 2,528 real commercial-advice queries and analysed 1,536 responses involving physical products. Their central operational lesson is not that one service is universally better or worse. It is that the observed recommendation depends on the environment in which it is measured.
See also: company profile
Confident language can hide unstable rankings
Among responses that recommended products, first-person preference language appeared at very different rates across the tested systems. At the same time, the selected products often changed across repeat runs. For brands, this makes one-off screenshots a poor basis for estimating “AI visibility.”
A useful audit needs repeated prompts, fixed comparison criteria and a record of the service, interface, account state, language, location and date. Without those controls, a change in recommendation can be mistaken for a gain or loss in brand performance when it may reflect normal output variation.
Interfaces and APIs are separate measurement surfaces
The study also reported low overlap between the domains displayed by different consumer interfaces for the same query. Consumer products and their APIs showed different source patterns as well. This matters because many monitoring tools rely on API access while customers make decisions in the public interface.
API data can still support scalable testing, but it should not be treated as a complete proxy. A defensible programme separates interface observation from API observation and reports the limitations of each.
Brands need an evidence-based AI discovery dashboard
The practical measurement unit should be an observable recommendation under a specified setup. Teams can record whether the brand was named, which product appeared, whether “best” framing was used, which links were visible and whether those destinations supported the claims made in the answer.
Repeatability is essential. Several runs of the same prompt reveal whether a brand appears consistently or only occasionally. Category-level prompts should also be distinguished from branded prompts, because they measure different stages of demand.
Do not convert a preprint into a market verdict
The work is a preprint and reflects selected systems, physical-product prompts and a defined testing period. It does not prove paid influence, systematic commercial bias or consumer conversion. Those questions require additional evidence, including longitudinal tests and behavioural outcomes.
The strategic conclusion is narrower and more useful: AI recommendation monitoring must be designed like an experiment. Brands that report a single answer as a stable market position risk optimizing against noise. Teams that measure repetitions, conditions, sources and change over time can build a more credible view of AI-mediated commerce.
