Our ongoing research points to specific content changes that increase AI-generated recommendations. In this dataset, the extent of this increase depended on several factors, including a brand's current visibility in a specific model and the context surrounding the brand's product.
The data shows that specific copy changes increased Gemini's generated recommendations from 83.3% to 100% and Claude's from 83.3% to 91.7%. ChatGPT-5.6 and GPT-4.1 both grew from 91.7% to 100% based on the two recommendations.

Brands invest heavily in rewriting product descriptions specifically to optimize them for LLM recommendations. Our research indicates that changing a product description adds little new value if a model already has broad awareness of a product. Each ChatGPT model picked the CeraVe-style moisturizer in 11 out of 12 comparisons before we changed the copy.
Most recommendations in Claude and Gemini occurred after changing the copy and were against the strong competitor products. In contrast, none of the recommendations changed in either comparison involving the medium or weak competitors.

However, the largest increase in recommendations came with a caveat: all six Gemini recommendation increases occurred when the CeraVe-style product was ranked second. Therefore, product rank, not a copy change, appears to have been at least partially responsible for the results, and additional testing would be needed to separate it from the impact of the copy change.
How This Data Can Help Beauty Brands Increase AI Recommendations
Before investing in a significant Geographic Optimization (GEO) or content rewrite initiative, establish a baseline evaluation of your brand's current recommendations and gaps. Then make small copy changes to test against competitors and identify relevant broad prompts where your brand can improve.
Investment rationale strengthens when you can show how a specific content change can repeatedly turn losing AI-generated recommendations into winning ones. Conducting a site-wide rewrite without evidence may create unnecessary expense with limited potential for improvement.
Mention copy:
“Sensitive skin in winter is a common seasonal skincare concern.”
Connection copy:
“This product is suitable for sensitive skin in winter.”
The first line in the product description stated the user’s need. The second line explained how the product met that need. All other lines (ingredients, etc.) were identical in both versions.
Gemini and Claude tested six competitors in two sets. Two were direct competitors with CeraVe: a fragrance-free moisturizer with ceramides, niacinamide, and glycerin; and another fragrance-free moisturizer with ceramides, hyaluronic acid, and amino acids.
Four brands were moderately competitive: a fragrance-free water gel with hyaluronic acid and glycerin; a fragrance-free moisturizer with petrolatum, glycerin, and sorbitol; a gel cream with hyaluronic acid, peptides, and squalane; and a lightweight make-up priming lotion with glycerin, panthenol, and algae extract. We ran each competitor twice against the CeraVe-style product, giving 12 head-to-head comparisons in each data set.
What content changes were worth making?
The answer depends on which model and test format. The table below provides the raw numbers from the analysis, excluding averages.
| Model/test | Need mentioned | Product connected to need | Recommendation results | Difference | p-value |
|---|---|---|---|---|---|
Gemini 3.5 Flash/reason + choice | 30/36 (83.3%) | 36/36 (100%) | 6/36 toward; 0 away | +16.7 pts | 0.03125 |
Claude Sonnet 4.6 / reason + choice | 30/36 (83.3%) | 33/36 (91.7%) | 3/36 toward; 0 away | +8.4 pts | 0.25 |
GPT-5.6 / letter only | 11/12 (91.7%) | 12/12 (100%) | 1/12 toward; 0 away | +8.3 pts | 1.0 |
GPT-4.1 / letter only | 11/12 (91.7%) | 12/12 (100%) | 1/12 toward; 0 away | +8.3 pts | 1.0 |
The data provide evidence that "user need relationships enhance AI recommendations": changing the product-to-need relationship into an explicit form altered AI recommendations, but the magnitude and reliability of the effects depended on the model used.
Under what conditions are content rewrites likely to have a limited return on investment?
We found that a rewrite is likely to have a low return on investment when the brand is recommended in most relevant comparisons. GPT-5.6 and GPT-4.1 each selected the CeraVe-style moisturizer in 11 out of 12 comparisons before adding the copy changes. After adding the copy changes, every model selected it in all 12 comparisons.
Each model had one losing recommendation left to change. Adding the new copy reversed that single loss on each model, but the starting rate (91.7%) left little room for growth.
While the test provided similar descriptions of the product's characteristics in both versions (cream, ceramides, hyaluronic acid, and glycerin; fragrance-free), we can’t determine whether any of those factors drove the original 11 recommendations because the ChatGPT runs only reported letters.
| Evaluation | Results | Implication |
|---|---|---|
How did the model answer? | Gemini baseline: 58.3% in the letter-only test vs. 83.3% in the reason-format test. Claude moved from 66.7%•50 % in the letter-only test, but from 83.3%→91.7% with a reason. | Answer instructions changed the starting rate and, for Claude, the direction of the measured effect. |
What defines a successful result? | Success meant selecting the CeraVe-style product. The study did not measure citations or brand mentions. | A recommendation, citation, and mention are different outcomes and should not be merged into one score. |
Was product order controlled? | Before the connection sentence, Gemini chose the product 18/18 times when it appeared first and 12/18 when it appeared second. Afterward, it chose it 18/18 in both positions. | All six Gemini switches occurred when the product appeared second, so the 16.7-point lift cannot be attributed to copy alone without another test. |
The explanation text changed what Gemini recommended. Before the connection (in 18 out of 33 explanations), Gemini used "Winter" or "Sensitive"; afterward (in 15 out of 33 explanations), when selecting products, it went up from 83.3% to 100%. Explanations and choices are two separate metrics.
Before you scale an AI search program, what kind of evidence should a pilot program produce?
A search program using an artificial intelligence search algorithm should tie a particular investment to specific changes in the results that matter to your brand.
This evidence supports an iterative process for content optimizations for specific LLMs: establish a model-level baseline, identify priority recommendations, make controlled changes, run multiple comparisons, and scale if the improvement is consistent.
Use measurable improvements as the foundation for an AI search program. One objective for brands is to show how a specific change improved a specific AI outcome enough to justify investing and scaling.
If your brand is exploring an investment in an AI search program, 5W helps identify which recommendations your brand can improve, develop model-level baselines, and pinpoint which content or media program modifications will drive the desired outcomes before spending money on larger programs. Start a conversation with 5W today to identify measurable AI-visibility gaps and learn what changes are worth testing to get started.
Methodology
Raw data files are available upon request.
1. Letter-only test (specified in run_sheet_96.csv → results in run_sheet_96_v2.csv [96 rows]; run_sheet_chatgpt_fix.csv [48 rows]). Gemini, Claude, GPT-5.6, GPT-4.1, and Perplexity were each instructed to respond with an a/B single letter to 12 head-to-head cells (2 direct/strong competitors and 4 moderate competitors, each with two copy conditions); no reasoning was captured.
The original chat_gpt rows in run_sheet_96_v2.csv failed (response = error: task status 40501 invalid field: ‘this…’, 0/24 valid target picks)—this is the “temperature” rejection referenced in the flags, confirmed in the data.
Perplexity’s 24 rows are all errors: task status 40501 invalid field: ‘web 'search'—confirmed 0/24 usable.
The run_sheet_chatgpt_fix.csv re-run demonstrates that GPT-5.6-terra and GPT-4.1 each had 23/24 valid target picks (11/12 in c2; 12/12 in c3), aligning with the reported 91.7% → 100%
2. Reason format test (run_sheet_reason_format_144.csv, 144 rows)—Gemini and Claude, each providing a short reason and a product code, 3 reps per cell/condition, at both product positions (first/second). Verified from the data:
Gemini: 83.3% (c2) → 100.0% (c3).
Claude: 83.3% (c2) → 91.7% (c3).
Gemini’s position breakdown: c2 = 100% (target first) v.s. 66.7% (target second); c3 = 100% in both positions—the entire 33-point gap disappears under c3, and this area is also the exact strata where all of Gemini’s movement occurs.
Claude’s position breakdown: c2 = 83.3% in both positions (no gap); c3 = 83.3% (first) v.s. 100% (second).
The post’s headline figures for Gemini/Claude (83.3% → 100%; 83.3% → 91.7%) come from the reason format test, while GPT-5.6/GPT-4.1's figures (91.7% → 100%) are based on the letter-only test. These are not equivalent experimental designs—different prompt structure, number of trials, and (for Gemini/Claude) three reps per cell versus one for GPT models—collapsed into a single comparative table as if they were identical.
Limitations (revised and re-verified)
1. Position is a confirmed source of bias for Gemini—it was never a questionable residual issue. Of Gemini's six “toward-target” flips, each one happened in a second-position cell—where the position gap had suppressed the target pick under c2 (66.7% → 100%). The data can’t distinguish between a copy effect and a rank effect for Gemini’s result—the cell type that moved was also the cell type where either explanation could remove a pre-existing bias.
2. Claude has the opposite directional movement than expected, and the result is a genuine, verified contradiction—not just a footnote curiosity. Letter only: Claude decreased from 66.7% to 50% (3 away; 1 toward; net negative), including the weak-tier flip observed in either test. Reason format: Claude increased from 83.3% to 91.7% (3 toward; 0 away). Same models and same copy variants—opposite direction—confirmed in both run_sheet_96_v2.csv/run_sheet_chatgpt_fix.csv and run_sheet_reason_format_144.csv.
3. There is no agreement across models on which cells change. In the reason format test, Gemini and Claude share one flip cell (R2-second) out of twelve; Gemini’s other flip (R1-second) is unique to it. In the letter-only test, none of the four cells flipped toward the target more than once. The results weakened any general mechanism claim for “connecting product to need helps,” making it more likely to be model-specific noise.
4. The stated reasoning-text mechanism doesn’t support itself and moves in opposite directions per engine—confirmed in the 144 row file. Gemini’s own mention of need in its own stated reasoning fell from C2 to C3 (18/36 → 15/36), the exact opposite of what “now the model is reasoning about the need” would predict. Claude’s rose (6/36 → 10/36 by counts referenced in the results file). The two models disagree on the metric meant to explain the effect.
5. The sample size for the “significant” result is small and concentrated. Gemini’s p = .03125 is based on 6 discordant trials, all of which fell within just two unique combinations of cell and position, rerun three times each—not six independent cells. GPT-5.6 and GPT-4.1's “effects” are single flipped rows each (p = 1.0—explicitly non-significant), and unlike Gemini/Claude, they were never rerun to check stability.
6. Two models were unable to be tested as specified due to technical issues. ChatGPT’s original run failed (0/24 valid) prior to the temperature-field fix; Perplexity failed and was never rerun; therefore, clean condition data does not exist for it in any file.
7. Baseline visibility is sensitive to answer format, not just copy changes. Gemini’s c2 rate is 58.3% letter vs. 83.3% reason format—a material difference between baselines dependent upon how the model is asked to respond.
8. Scope is narrow. All data came from one category (moisturizers) vs. a fixed set of six competitors and one pair of copy variants. None of the four run sheets tested a different product, a different need statement, or more than two copy variants, so the results shouldn’t be read as general findings regarding “content changes” beyond this specific setup




