Market access & commercial

Consumer concept and A/B testing

The inference problems in consumer research: hypothetical bias in stated preference, non-probability panels, the sample size an A/B test actually needs, and why looking at a running test breaks its error rate.

Consumer research and a trained sensory panel are often run by the same agency and confused for one another, but they answer different questions and are limited by different things. A discrimination panel asks whether a difference is perceptible at all, and its statistics are the clean binomial arithmetic set out on the sensory panels page. Consumer research asks whether people will choose the product, and that quantity is not perceptual. It is a prediction about future behaviour made from present statements, and every methodological difficulty in the field comes from that substitution.

Stated preference is not revealed preference

The gap between what respondents say they would pay and what they pay is called hypothetical bias, and it is not noise: it is directional. Willingness-to-pay elicited in surveys routinely exceeds willingness-to-pay observed at a till, because a stated answer costs nothing and a purchase costs money.

The bias is worst for socially approved attributes, which is precisely the territory of bio-based, recycled, cruelty-free and low-carbon claims. Respondents are answering in the presence of a researcher and of their own self-image, so the same attribute that motivates the product also inflates the response to it. This is the documented shape of the attitude–behaviour gap in sustainable consumption, and no amount of sample size corrects it, because it is bias rather than variance. Partial remedies exist — incentive-compatible designs where a choice is occasionally binding, cheap-talk scripts, discrete choice experiments that force trade-offs against price instead of rating an attribute in isolation — and all of them work by making the answer cost something.

The sample is not the population

Online panels are non-probability samples: respondents opted in and were routed to the study. Quota sampling fixes the marginal distribution of age, region and income, but it cannot fix selection on unobserved variables, and interest in the product category is exactly such a variable. The confidence interval printed on the report is computed as if the sample were random, so it describes sampling error only, and understates total error by an unknown amount. Focus groups have a harsher version of the same problem: participants hear each other, so responses inside a group are not independent observations, and the effective unit of analysis is the group, not the person in it.

What an A/B test can and cannot settle

Randomised assignment does buy causal identification, within the tested population and period. The costs are arithmetic. The sample required to detect an effect scales roughly with the inverse square of that effect, so halving the smallest difference worth finding quadruples the traffic — which is why small absolute lifts on low-conversion pages are frequently undetectable in any realistic time.

Two failure modes are routine. Testing many variants inflates the chance that at least one appears significant, unless the threshold is adjusted. And monitoring a running test and stopping when it crosses significance invalidates the fixed-horizon p-value entirely, because repeated looks are repeated chances to cross by luck; either fix the horizon in advance, or use a sequential procedure designed for continuous monitoring.

Finally, a claim that tests well may still be unlawful. Consumer preference is not evidence of truth, and in the EU the substantiation of a health claim is a separate matter decided on the evidence EFSA requires.

Last updated: