Columbia Study Says Retail AI Detects but Does Not Flag Origin Conflicts

A Columbia Law School center reported on July 30 that Amazon's Alexa for Shopping and Walmart's Sparky identified contradictions in some "Made in USA" listings during its tests, yet did not flag those claims for shoppers. The report relies on selected chatbot interactions and listing checks, so its results are evidence from the researchers' investigation rather than a platform-wide audit of either retailer's enforcement systems.
A Columbia Law School center reported on July 30 that Amazon's Alexa for Shopping and Walmart's Sparky identified contradictions in some "Made in USA" product listings during its tests but did not flag those claims for shoppers. The finding comes from the researchers' documented interactions with the chatbots; it is not a measured failure rate across either marketplace.
What the researchers tested
Erie Meyer and Zachary Harris of Columbia Law School's Center for Law and the Economy examined whether the two shopping assistants could surface products, assess suspicious origin claims, and respond differently to equivalent queries about U.S.-made and foreign-made goods. When a chatbot refused a query, the researchers changed the wording to test whether the information was unavailable or a particular phrase was triggering the refusal.
The report documents an Amazon comparison chart that described six T-shirts as "Made in USA" even though the linked product pages described them as imported. When prompted about the conflict, Alexa for Shopping identified the discrepancy. On Walmart, the researchers asked Sparky to compare products and rank how suspicious their U.S.-origin claims appeared using listing details and manufacturing-footprint information.
The researchers also reported that Alexa for Shopping refused a request for "Made in USA" fly-fishing reels but returned U.S.-made recommendations after they changed the wording to "in USA." They interpreted that result as evidence of a phrase-specific guardrail rather than a genuine lack of data.
What the evidence does and does not establish
The screenshots and listing checks support the narrower conclusion that the tested assistants could reason over some origin-data conflicts while the listings remained unflagged. They do not establish a platform-wide prevalence rate, measure every marketplace control, or independently prove why Amazon or Walmart configured their systems as they did.
That distinction matters because the chatbots also generated explanations about business incentives and enforcement pressure. Those answers are model outputs reported by the researchers, not verified statements of internal company policy or evidence of executive intent. The report says neither company responded to the authors' requests for technical corrections; The American Prospect and Quartz covered the release but did not independently reproduce the tests.
The regulatory context
The Federal Trade Commission told Amazon and Walmart in July 2025 that third-party sellers' unqualified U.S.-origin claims can violate the FTC Act or the Made in USA Labeling Rule unless a product is "all or virtually all" made domestically. The FTC asked the marketplaces to monitor and address misleading seller claims, while its Amazon letter expressly said it was not an assessment that Amazon itself had violated the law.
For ML and trust-and-safety teams, the practical lesson is that detection is only one layer of enforcement. A production system also needs reliable provenance, contradiction thresholds, human escalation, remediation rules, and audit logs that connect a model's signal to an accountable marketplace action.
Key Points
- 1Columbia researchers documented selected cases in which Alexa for Shopping and Sparky identified origin-data conflicts without flagging the listings for shoppers.
- 2The tests show a capability gap between detection and action, but they are not a platform-wide benchmark and chatbot explanations do not prove company policy or intent.
- 3The FTC's 2025 letters set the legal backdrop: unqualified Made in USA claims generally require products to be all or virtually all domestically made.
Scoring Rationale
The report connects shopping-assistant behavior at two major marketplaces to product-origin disclosure and consumer-protection controls. It is useful to teams building conversational commerce and marketplace enforcement, but the evidence is based on selected researcher tests rather than a platform-wide benchmark.
Sources
Primary source and supporting public references used for this report.
Practice with real Ad Tech data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Ad Tech problems
