The answers are surprising. New 5W research helps brands preparing for Black Friday answer this question.
We tested how Black Friday discounts impact AI product recommendations using five anonymous noise-canceling headphones in each trial:
We used the following prompt modeled after a user looking for a mid-to-high-end gift: "I'm Black Friday shopping and looking for noise-canceling headphones to commute to work and for travel. " Which products do you recommend?"
We set a price cap of $300. All product specs stayed static, and the only factor that varied between trials was the Black Friday discount offered on individual products (and therefore their sale price).
Products A, C, and D were tested at three Black Friday discount levels (10%, 25%, and 40%) ($270, $225, and $180). The remaining four products were sold at full price ($300) and served as the sales baseline. Each model tested a random order of products to eliminate any positional bias that could be attributed to the discount.
We ran each condition five times using these three models: Claude Sonnet 5 (Anthropic), gpt-5.6 sol (OpenAI), and Gemini 3.5 Flash (Google). All observations provided one product recommendation. None of the models provided brand names, reviews, ratings, retailers, or any information regarding which product was discounted.
How Did a Black Friday Discount Take Product C From 0% to 53%?
Product c was never recommended at full price or at a discount of 10%. However, at a 25% discount, Product C was chosen in 2 of 15 trials. At a 40% discount, Product C was chosen in 8 of 15 trials (53.33%).

The number of trials is almost as important as the percentages, explaining why Product C was more likely to be recommended at a deeper discount. There were enough trials (n=15) to estimate confidence intervals around the number of trials in which Product C was recommended at a 25% discount (range = 4.34 - 21.66%) and at a 40% discount (range = 29.84 - 74.16%). These ranges don’t overlap with each other or with the full-price scenario range.
As expected, the models' output and referenced price, savings, the sale, etc., more often in their rationales when Product A was discounted than when Product C was discounted.
In fact, every model wrote rationales that referenced price/savings/etc. in each baseline trial and continued to write rationales at 10%, 25%, and 40% off for Product A. However, when Product A was not discounted, none of the models referenced price/savings/etc. in their rationales.
In contrast, the discount did not appear in the outcome. Product C was never recommended when Product A was discounted by 10%.
What Did a 25% Discount Do to Product C's Recommendations?
When Product C was discounted by 25%, it received 2 recommendations in 15 trials. One came from GPT-5.6 Sol (OpenAI), and the other came from Claude Sonnet 5 (Anthropic).
We identified 25% as the first tested discount amount where Product C's recommendation rate rose above its baseline, but we don't think 25% is a threshold. With only 2 of 15 recommendations in the dataset, we can't draw conclusions from this limited evidence. Additionally, 25% is just one of four discount amounts we tested and likely represents one point on a curve.
What Made Product C the Top AI Pick at 40% Off ($180)?
At $180 (40% off), Product C was recommended in 8 of 15 trials (a 73% increase relative to its full-price baseline at zero).
This number is the main finding from this study. A product that no model recommended at full price ($300) became the most recommended product at a discounted price ($180), even though all specifications, competitors, and the user prompt remained identical.
Why Didn't a 40% Discount Help Product D Get Recommended?
Similar to Product C, Product D started with zero recommendations at full price ($300) and received the same 10%, 25%, and 40% discounts, but it still finished with no model recommendations.

Fifty-six trials later (60), including 15 trials at $180 and a $120 price drop that made Product C the most frequently recommended, failed to impact Product D.
Interestingly, Product D has longer battery life than any other product (50 hours), but it also has lower-rated comfort (3/5), is heavier (280 g), and has a mediocre call microphone. In contrast, Product C has higher-rated call quality (the highest among the five), lighter weight, and an average comfort rating (4/5). The study did not determine which specific characteristics caused the difference in recommendations and remained focused on price cuts.
Product A represented the opposite scenario. Each model recommended it in all baseline trials and also recommended it in 15 trials at 10% off, 25% off, and 40% off. Thus, Product A was selected in every trial before and after the discount, maintaining a 100% recommendation rate throughout. The discount clearly affected the product's visibility, as each product was equally visible at full price and no additional products competed for attention after the discount.
Two additional findings are also worth noting. Product E was never recommended in any trial. Product B was selected twice; both instances occurred at 25% off: once when testing for Product C and once when testing for Product D, so one recommendation per call is likely noise.
How Did GPT-5.6, Claude Sonnet 5, and Gemini Flash Respond to Discounts?
Of the 8 selections for Product C at 40% off, 6 came from two models (3 from Sol and 3 from Sonnet).

GPT-5.6 recommended Product C in every trial at 40% off. Claude Sonnet chose it in 3 of 5 trials. Gemini Flash never recommended it regardless of price level.
Overall, the three models disagreed on which product to select in 19 of 50 comparable trials (i.e., 38%). Of those 19 differentiations, 18 occurred within the trials testing Product C at either 25% or 40% off. Within those five Product C trials at 40% off, each model recommended it differently.
Before the discount triggered recommendations, the models agreed almost unanimously on a single product recommendation.
Can an AI Model Weigh a Discount and Still Not Recommend the Product?
Product D is the best example of this scenario. At 10% off, 7 of 15 model explanations referenced price, savings, or a deal but excluded Product D from its recommendations. At 40% off, 4 out of 15 explanations referenced price and continued to exclude Product D from their recommendations.
This finding examines keyword matches in the model's answers: did words like "discount," "sale," "savings," "deal," "value," or "the sale price" appear? The answer doesn't explain the model's reasoning, but it does show that a discount can appear in a model’s evaluation process without appearing in its recommendations. Future 5W research will help determine the cause for that exclusion.
These findings help show why tracking AI recommendations is critical for brands. Recommending a product discount is not the same as recommending the product.
How Can Brands Use This AI Recommendation Research?
In 150 trials, a 10% price decrease had no effect. A 25% price discount resulted in 2 recommendations out of 15 for one of the three products. With the discount raised to 40%, that product went from zero recommendations to 8 out of 15. An additional product was never recommended regardless of discount. The third product had maintained consistent recommendations (15 out of 15) since it was originally recommended at full price; even a 40% discount could not increase that number. Only two of the three models responded to a 40% discount, and one model never responded to any discount.
We found that price could influence a product's ranking through a generative recommendation. However, the extent to which price influences the ranking and the degree of that influence vary depending on the specific products being compared and the version of the model providing the response.
Brands looking to promote items on Black Friday should test their products at specific price points using a preferred AI model before and after the promotion. Identifying product mentions alone won’t provide much insight, and success should be measured by recommendations.
Conclusion
Discounts can move an AI model from excluding a product to recommending it, but an identical discount left the recommendation status of a product excluded at baseline unchanged. To learn which side a product sits on, brands should test a specific price point across several models and measure recommendations, not mentions.
Methodology
5W tested a single, common user prompt to learn if a Black Friday discount changes which product a generative model recommends. Five anonymized noise-canceling headphones (Products A through E) remained constant across every run; only the discount and sale price changed, one product at a time. The baseline price was $300 in every condition, discounted to $270 (10% off), $225 (25% off), and $180 (40% off). The words “Black Friday” appeared for each product in all conditions, including the shared 0% baseline.
The same prompt went to three models on every run: Anthropic's Claude-sonnet-5, OpenAI's GPT-5.6-sol, and Google's Gemini-3.8-flash. The study recorded 150 observations, 15 at baseline and 135 across the discount conditions, with 5 trials per condition per model. We randomized product display order from a stored seed to remove position effects, and no run included brand names, reviews, ratings, retailers, scarcity, or shipping cues. We report selection rates with underlying counts and Wilson 95% confidence intervals, and we express discount effects as percentage-point changes rather than percent changes. A discount “mention” is a keyword match in a model's short rationale, not a measure or indication of model reasoning. All 150 planned observations produced a valid parsed response, and a replication on gemini-3.5-flash matched gemini-3.8-flash on all 50 comparable trials (100%).
Limitations
These results describe the tested setup, not general generative model behavior. The results cover five products, one prompt, three models, and a shared 0% baseline with three discount levels (10%, 25%, and 40%). Different product sets, requests, model versions, or price ranges produced different outcomes.
The study measured which product each model selected, not why. It did not isolate which product attributes drove a selection, and the discount “mention” rate is a text signal only, not evidence of internal reasoning. Provider reasoning and sampling settings varied across the three models and were captured inconsistently, so read cross-model differences with that limit in mind. Observed thresholds mark the lowest tested discount that beat a product's baseline by at least 10 percentage points; they are observed points for these conditions, not a universal or optimal discount.





