Researchers propose bias correction for demand models using unstructured data

Timothy Christensen and Giovanni Compiani have released a revised working paper that adds a post-estimation bias correction and proxy-quality diagnostics to consumer-demand models built from text, images, and other unstructured data. In an experiment with 9,265 participants choosing among 10 ebooks, their preferred specification raised the hit rate for predicting the closest substitute from 40% to 70%. The method is intended to make pricing, assortment, and merger counterfactuals less sensitive to imperfect ML-derived product attributes.
Timothy Christensen of Yale and Giovanni Compiani of the University of Chicago have developed a statistical toolkit for a problem that appears when economists use machine-learning representations of product text or images in consumer-demand models: the embeddings may omit or distort attributes that actually drive substitution. The revised working paper introduces a post-estimation bias correction and diagnostics for judging whether a proxy is adequate before analysts rely on a counterfactual result.
The measurement problem
Demand models are often used to estimate what consumers would do after a price change, product removal, merger, or assortment decision. Text descriptions, reviews, and images can enrich those models, but an embedding is still a proxy for the latent qualities consumers value. Treating that proxy as perfectly measured can bias both parameter estimates and the resulting counterfactuals.
The authors' correction maps discrepancies between the fitted model and observed data into an adjustment to the counterfactual estimate. They also propose two diagnostic tests: one evaluates whether the estimated proxy-driven representation is close enough to the latent target for the correction to be credible, and the other checks whether the proxy has sufficient dimensionality. The paper says the method works with market-level or individual data, supports proxies produced by fine-tuned models, and adds relatively little computation after the underlying demand model is fitted.
What the experiment found
The empirical application uses data from 9,265 participants who chose among 10 ebooks after seeing prices, standard attributes, cover images, descriptions, and reviews. After each participant's first choice was removed, the experiment recorded a second choice, creating a direct test of whether a model could identify the closest substitute.
For the authors' preferred specification, the bias correction increased the closest-substitute hit rate from 40% to 70%. Across the review-based specifications, the paper reports improvements from 40% to 60%-70%, while a random choice among the remaining books would achieve about 11%. The result is an application-specific validation in a working paper, not proof that the same gain will carry to every market or embedding pipeline.
Why it matters
For data scientists working on pricing, recommendations, assortment, or merger analysis, the practical lesson is that predictive fit on observed choices does not guarantee reliable counterfactuals. Proxy diagnostics and sensitivity checks should sit between embedding generation and a business or policy decision. The proposed method gives applied teams a way to quantify part of that measurement risk while retaining familiar demand-model workflows.
Key Points
- 1The working paper adds a post-estimation correction for counterfactual bias caused by imperfect ML-derived product attributes.
- 2In a 9,265-participant ebook experiment, the preferred specification's closest-substitute hit rate increased from 40% to 70% after correction.
- 3The authors also propose diagnostics for proxy quality and dimensionality, but the empirical gain remains application-specific and has not been independently replicated.
Scoring Rationale
The paper offers a concrete correction and diagnostics for counterfactual models that use ML-derived text and image proxies, with a sizeable application result and direct relevance to pricing and recommendation work. Impact is tempered because it is a working paper and the 40%-to-70% result comes from one ebook experiment rather than independent replication.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems


