Executive Summary
Consumers’ ability to discover information through artificial intelligence (AI) has created a new opportunity for wineries. While users can find a specific wine based on style, seasonality, price point, flavor profile, and/or reputation, they lose the benefit of AI recommendations if the data in available content cannot link the wine to the specific details in their prompt (query).
In our controlled test, we showed quantifiable differences between adding additional product descriptions and adding more relevant product context. In **71.1% of control trials**, the target wine was chosen; in **72.2% of trials, using only a neutral copy**; and in **76.7% of trials** when "summer dinner party" was shown but did not indicate that the wine would be well-suited to this use. When the product information indicated that the wine would be well-suited to a summer dinner party, the trial success rate was **100% over 90 blinded trials**.
This means that rather than requiring longer product pages, winery executives should look at how their websites, retailer content, media exposure, and all other forms of product content explain **why a particular wine is a good option, what type of consumer it best fits, and which events or uses it is best suited for**. We found that connecting the product to the consumer's need improved results far more than adding neutral content or restating the request without building that relationship.
Our Findings: What Information Increased AI Wine Recommendations
Recommendations increased most when the wine was specifically linked to users' needs.
In 90 blinded trials, the target wine was selected 64 times based on website product descriptions, 65 times based on neutral copy, 69 times when the term "summer dinner party" was mentioned without a link to the product, and 90 times when the description indicated the wine was suitable for a summer dinner party.
Consequently, wineries should identify the user request or occasion (in this study, a dinner party), how their products support it, and communicate that relationship rather than leaving the model to infer the association. Wineries should add content that addresses common user needs rather than creating excess content without providing contextually relevant uses.
The study shows that the specific user request produced significantly fewer changes than establishing the connection between the user terms and the product.
Gemini improved from 11 of 30 recommendations in control trials to 12 of 30 in trials where the term "summer dinner party" was referenced without indicating the wine was suited to that occasion. Gemini reached 30 of 30 when the relationship between the product (wine) and the occasion (dinner) was specified. Word choice related to the occasion resulted in favorable changes (21 pairs) over reference language.
Overall, 21 trials showed that adding relationship language led to a positive outcome (the target wine was recommended), while none resulted in a negative outcome.
These results suggest that wineries should examine whether their product descriptions and content address specific user needs over detailed ingredients or other features.
Blinded Test Results: Recommendation Rates Increased From 71.1% to 100%
| Condition | Number selected | Recommendation rate |
|---|---|---|
Control | 64 of 90 | 71.1% |
Neutral Copy | 65 of 90 | 72.2% |
Request Language Only — Added "summer dinner party," but did not indicate that the wine is well-suited for that occasion. | 69 of 90 | 76.7% |
Product-to-Use Relationship — Added "this [product] is suitable for a summer dinner party." | 90 of 90 | 100% |
Why AI Models Often Miss a Wine That Fits A User’s Request
Conversational queries have become common as users interact with AI models.
For this study, we asked ChatGPT to provide "a good rosé for a summer dinner party" rather than searching for "best rosé wines." This request highlights three key elements that shape the model's recommendation.
In this scenario, the user is looking for a rosé wine suited to a particular season (summer) and occasion (dinner party).
We began researching the topic after manually running multiple searches in ChatGPT, Claude, and Gemini with the request, "recommend a good rosé for a summer dinner party."
Surprisingly, Summer in a Bottle was omitted from results across each model even though it matched two of the three attributes identified in the query. Summer in a Bottle is a well-regarded and popular rosé and references "summer" in its title. The remaining attribute was whether or not the wine is well-suited to the identified occasion ("dinner party"). Therefore, the experiment focused on determining if describing the relationship between the wine and the user's occasion influenced the model's decision-making process.
LLMs identified "Summer in a Bottle” for “Rosé and “Summer” but not "Dinner Party"
As stated above, the user's initial request identified three distinct elements: product type, time/season, and occasion.
Summer in a Bottle matches two of these elements: product type and time/season. The model needed to connect Summer in a Bottle to an occasion to satisfy the third element.

The experiment did not measure how often including "rosé" or "summer" would improve model performance. Instead, it measured how models performed differently when content (product display pages, media, on-site content, etc.) linked the wine to a user's situation.
AI models understand what a product says on a page and a user's needs, but they don't consistently associate products with those needs unless the available information connects them.
How 5WPR Tested What Drives AI Wine Recommendations
Our main analysis examined 360 model responses through blinded experiments involving ChatGPT, Claude, and Gemini. The analysis considered eight competing wines. The models received static data on each wine rather than dynamically retrieving it via live web searches.
We removed all brand/product names during blinded modes, so prior knowledge of Summer in a Bottle or Wölffer Estate would not explain any results.
Each model completed thirty trials under four different conditions (i.e., 120 trials per model), with ninety trials per condition per model.
All models received the same user request: "recommend a good rosé for a summer dinner party."
Each model's underlying information remained consistent. The only difference between conditions was one sentence about evidence for each target wine description.
The Four Content Conditions Tested Across ChatGPT, Claude, and Gemini
Control: no additional sentence was added to either model input.
Neutral Copy: added, "[Product] should be served chilled when poured." While this text added similar-length copy, it did not mention a "summer dinner party."
Request Language Only: added, "Summer dinner party traditions vary by region and season." Although this language added the phrase "summer dinner party," it did not state that the wine was suitable for this occasion.
Product-to-Use Relationship: added, "[Product] is suitable for a summer dinner party." The last condition linked the target wine with the user prompt (as defined within their request).

Candidate order/blind identity was kept constant between conditions within each pair of trials. Search functionality and web-searching capabilities were disabled during experimentation. The experiment collected 720 responses (blinded + branded), with no API failures or unparseable responses.
Neutral Copy Produced Almost No Increase in Recommendations
The neutral copy condition evaluated whether an additional sentence would improve performance by increasing attention/more content to the target description.
Claude scored 25/30 in both the control and neutral copy conditions. Gemini also scored similarly at 11/30 in both conditions. ChatGPT rose slightly from 28/30 to 29/30.
As shown above, neutral copy created little change overall, moving from a 64/90 = .711 overall recommendation rate for standard information to 65/90 = .722 overall.
Recommendation rate for neutral copy. The paired t-test comparing these two conditions showed p = 1.0.
This study also shows that adding additional serving details (like chilled vs. room temperature) did not create the same gain as adding a sentence that specifically addressed the occasion-related part of the user's request.
Using “Summer Dinner Party” Without Linking It to the Wine Produced a Smaller Increase In Recommendations
The next condition included the same phrase, "summer dinner party," in product information without stating that the wine suited the occasion (dinner party).
Claude rose from 25/30 control recommendations to 27/30. Similarly, Gemini rose from 11/30 to 12/30. ChatGPT rose from 28/30 to 30/30.
Overall, the recommendation rose from .711 to .767.
While neither the control nor the request language-only copy condition reached statistical significance (p = .125), the findings show a separation between phrase usage and intended meaning.
Although the user’s terms were present in both conditions, the models had to infer whether they understood which product best suited the occasion.
Linking the Wine to the Consumer’s Occasion Increased Recommendations to 100%
The final condition modified one piece of information relating to the “Summer In a Bottle” brand: [Product] is suitable for a summer dinner party."
In all three models, every blinded trial produced a recommendation for Woeffler. Thus, recommendation rates were 100 percent across all conditions for every model.

Gemini showed the most significant change among the three models.
Recommendations rose from 12/30 with Request Language Only copy to 30/30 with Product-to-Use Relationship copy—an increase of 60 percent.
Claude rose from 27/30 to 30/30. ChatGPT remained at 30/30 because its Request Language Only condition had already maximized recommendation possibilities.
Among all models, product-to-use sentences increased recommendations over Request Language Only by an average of .233.
Additionally, twenty-one instances showed that specifying relationships improved recommendations compared with using only user terminology in copy—twenty-one trials moved from failing to select the target wine to selecting the target wine, and zero trials moved from selecting the target wine to failing to select the target wine—statistically significant at p < .00001.
Gemini Showed the Largest Increase in AI Wine Recommendations
| Blinded condition | ChatGPT | Claude | Gemini |
|---|---|---|---|
Control | 93.33% | 83.33% | 36.67% |
Neutral Copy | 96.67% | 83.33% | 36.67% |
Request Language Only | 100% | 90% | 40% |
Product-to-Use Relationship | 100% | 100% | 100% |

These percentages represent model-specific performance; however, they show where larger differences existed. Specifically, Gemini performs better because it started from a lower baseline—i.e., only 40% of blinded trials previously chose Gemini as the best option—so it had room to grow.
When copying switched from referencing occasion to stating that the product is well-suited to that occasion, Gemini's recommendation rate jumped from 40% to 100%—an increase of 60%. The p-value associated with Gemini's comparison was .000008.
Claude also increased in this direction, although its .9 request language only rate-limited further growth by leaving three additional recommendations before it reached a 100% recommendation rate.
Because ChatGPT achieved a 100% recommendation rate before the product-to-use relationship copy was added, it established an upper bound for detecting additional increases in top-three recommendations for ChatGPT.
How the Blinded Results Tested Content Changes
This study conducted an additional set of examinations using actual product names visible to models (360 responses)—these branded results offer less clarity in isolating effects, primarily because some models showed large changes in recommendations when neutral sentences were added.
Claude offers the most obvious outcome. In blinded mode, when the name "Summer in a Bottle” was visible, Claude selected “Summer in a Bottle” one out of thirty times, or approximately 3.5% of trials.
However, when a neutral serving sentence was added and did not include any reference to a "summer dinner party," Claude selected "Summer in a Bottle" eighteen times (or 60%), even though the neutral sentence did not specify the relationship between “Summer in a Bottle" and "summer dinner party." This outcome did not hold in blinded mode. When brand names were obscured, Claude performed at levels similar to both the blinded control (twenty-five-of-thirty recommendations) and Nneutral Copy (also twenty-five-of-thirty recommendations) conditions.
Therefore, the blinded trial accurately tests the research question because increasing length did not produce the large movement observed under branded-mode conditions.
How Wineries Can Improve Their Content and AI Search Strategy
This experiment does not suggest that wineries need excessive product descriptions. Evidence against this conclusion exists in the neutral copy condition: Claude and Gemini produced no additional recommendations after adding neutral copy.
Rather than supporting longer descriptions, these findings point to specific content improvements.
Wineries often share helpful information about grape variety, region, year-vintage, flavor descriptors, acidity/sweetness levels, production methods, and optimal serving temperatures, but an AI model often cannot infer how that content communicates whether the bottle fits the occasion described in the user's recommendation inquiry/request.
As illustrated by the Summer in a Bottle test, communicating the relationship between product and use produces much larger variations in recommendations than communicating ingredients or other details used by customers requesting advice/recommendation.
Examples of other types of occasions a winery could describe include but aren’t limited to:
A wine for an outdoor wedding;
A red wine for steak;
A bottle to bring as a gift/host;
A sparkling wine for brunch;
A wine for holiday dinner; or
A rosé for seafood served outside.
Wineries should identify strong relationships supported by products and communicate them; the lesson from this experiment isn’t to invent new use cases but to document existing relationships supported by products that are absent from the product information provided to models.
What Wineries Should Take From the AI Recommendation Study
Even when a wine fits several components of a user request, it remains infrequently recommended unless available information explicitly links the product information to a user need.
To summarize the results: The control resulted in a seventy-one point eleven percent recommendation rate; a neutral copy resulted in a seventy-two point two percent recommendation rate; request language only resulted in a 76.7% recommendation rate. product-to-use relationship: Stating that the product is suitable for the users' requests resulted in a 100% recommendation rate across 90 blinded trials.
Convert Your Wine Portfolio into a Stronger Source for AI Recommendation Opportunities
Start by focusing on the wines within your portfolio that drive the greatest amount of revenue, distribution, and/or growth. Then identify the consumer requests each bottle can satisfactorily meet—e.g., **wine for an outdoor wedding, red wine for steak, host gift, champagne for brunch, wine for a holiday dinner,** etc.—and assess whether you currently communicate these relationships within your existing content.
**5WPR will review your wine portfolio relative to the consumer requests driving AI recommendations, identify gaps in product-to-use relationships, and help ensure your website and earned media content provide the right information for AI systems to easily locate and understand these connections.
Study Limitations
The study evaluates one aspect under controlled circumstances. The study cannot explain why Summer in a Bottle was excluded from previous live search queries; commercial models access different data sources and apply processing techniques unavailable in the controlled test environment.
Experimental recommendation rates do not reflect the estimated probability of recommendation success in real-world settings.
The experiment establishes no causal link between posting relationship language on winery websites and corresponding increases in real-time product recommendation performance; it presented direct evidence to models and did not evaluate impact on crawling/indexing/retrieval/source.
Findings are specific to one type of user request: eight competing wines, three fixed model versions, and tested product information; therefore, these results may not generalize across varying wines, requests, competitor products, or model versions.
Additionally, ChatGPT started near the upper limits of the measurements used in this experiment, limiting its ability to quantify additional increases in top-three recommendations for this model version.
Study Methodology:
5WPR designed this controlled experiment to test this question: Does connecting a wine to a consumer’s stated occasion affect whether a model recommends the wine?
The experiment was designed to isolate that effect from brand recognition, additional copy, keyword use, candidate position, and outside information.
Models and User Query
The study tested pinned versions of three language models:
OpenAI: GPT-4.1 (gpt-4.1-2025-04-14)
Anthropic: Claude Opus 4.5 (claude-opus-4-5-20251101)
Google: Gemini 3.5 Flash (gemini-3.5-flash)
The temperature was set to 0.0. Each model received the same consumer request:
“Recommend a good rosé for a summer dinner party.”
The models were instructed to use only the product information supplied during the experiment and recommend exactly three wines. Web search and outside product knowledge were excluded from the experimental task.
Brands
Each trial presented the same eight competing rosé wines:
Wölffer Estate — Summer in a Bottle
Château d'Esclans — Whispering Angel
Château Miraval — Miraval Côtes de Provence Rosé
Domaines Ott — By Ott Rosé
Château Minuty — M de Minuty Rosé
Gérard Bertrand — Côte des Roses Rosé
Commanderie de Peyrassol — Château Peyrassol Rosé
Mas de Gourgonnier — Rosé Tradition
The product evidence cards were researcher-written summaries based on product information from winery-owned sources. The underlying source pages were archived, and substantive details—such as region, grape varieties, and tasting descriptors—were checked against those records.
Conditions
Summer in a Bottle served as the target wine. The descriptions of the other seven wines remained unchanged throughout the experiment.
Four versions of the target information were tested:
Control: No additional sentence.
Neutral Copy: “[Product] should be served chilled when poured.”
Request Language Only: “Summer dinner party traditions vary by region and season.”
Product-to-Use Relationship: “[Product] is suitable for a summer dinner party.”
The neutral condition tested whether adding another sentence alone affected selection. The request-language condition tested whether including the words “summer dinner party” affected selection without connecting the phrase to the wine. The final condition tested the product-to-use relationship itself.
No treatment language was published on a winery website or presented to consumers as part of the experiment. The sentences were created solely for controlled testing.
Blinding, Randomization, and Trial Design
The completed experiment contained 720 successful model responses: 30 trials for each of four conditions, across three models, in both blinded and branded modes.
The 360 blinded responses formed the clearest test of the content effect. We replaced brand and product names with neutral labels such as “Item A” through “Item H,” reducing the influence of prior brand familiarity.
Candidate order and blinded aliases changed across trials using a deterministic seed-based process. All four conditions within the same paired trial used the same candidate order and aliases. The tested sentence was therefore the primary difference between the paired conditions.
A separate 360-response branded analysis retained the actual wine names. We analyzed the branded results separately because existing model associations with recognizable brands could affect selections.
We counted a target selection only when the model returned an approved name or corresponding blinded alias for Summer in a Bottle. Matching used deterministic rules that normalized capitalization, accents, and punctuation.
No language model graded, interpreted, or classified another model’s recommendations. Rule-based scoring prevented a second model judgment from becoming part of the measurement process.
Every completed call recorded the model, model version, condition, trial number, randomization seed, candidate order, target position, blinded alias when applicable, recommendation rank, raw recommendation, parsing status, token use, and other execution data. We also retained complete raw model responses.
Statistical Analysis and Research Transparency
We evaluated paired conditions using an exact McNemar test, which examines trials in which the same model selected the target under one condition but not the paired comparison condition. The design allowed us to measure whether changes consistently shifted selections in one direction while holding candidate order and aliases constant.
The research record also preserves unsuccessful or superseded experimental runs and documents why those runs were excluded from the reported findings. Excluded runs were not used to calculate the results presented in this study.
The experiment should be interpreted as a controlled test of how supplied product information affected recommendations under the tested conditions. The methodology does not treat experimental selection rates as forecasts of live AI search performance; the Study Limitations section distinguishes the controlled environment from real-world search.
The full study data set is available upon request, including trial-level results, treatment conditions, statistical outputs, source records, and raw model responses.





