Publishers that explicitly block AI crawlers in their robots.txt get cited by Google's AI Mode more than three times as often as the baseline rate, 51.9% of cited domains versus 15% across the general sample, according to a July 2026 baseline study of 10,894 domains by HasData. NYT, CNN, CNBC, and the BBC all block multiple AI crawlers in robots.txt. All four still turned up repeatedly as cited sources in the same study's AI Mode test.
That single finding should change how most technology PR teams build a media list. “AI-friendly” has become shorthand for “doesn't block AI crawlers,” and that shorthand is measuring the wrong thing.
Does Placing a Story in an “AI-Friendly” Publication Matter?
Checking whether a target publication blocks GPTBot in robots.txt tells a tech PR team almost nothing about whether a placement there will get cited. What predicts citation is the same thing that always predicted it: independent, credible, current journalism, not a technical flag in a file most readers never see.
A separate study of 4 million citations across 3,600 prompts, run by Citation Labs' XOFU tool and authored by BuzzStream's Vince Nero, found the relationship between crawler access and AI citation to be far weaker than most publishers assume, according to PPC Land's coverage of the March 2026 research. Blocking a crawler is largely a separate decision from whether a story ends up cited.
Why Do Blocked Outlets Still Get Cited?
Because the directive a publisher writes and the directive that actually gates an AI answer are often two different settings entirely. Google-Extended controls whether pages already sitting in Google's search index can ground an AI Overview or AI Mode answer. GPTBot, ClaudeBot, and CCBot control a completely separate question: whether that crawler can fetch pages for its own training set. A site can block all three training crawlers and still show up in Google's AI answers through the same search index it was never trying to leave, per HasData's research.
Roughly a third of top sites now run what amounts to a middle-path policy: blocking training-class bots like GPTBot and ClaudeBot while leaving answer-oriented bots like OAI-SearchBot and PerplexityBot untouched, according to Presenc AI's 2026 tracker of robots.txt adoption. A blanket “this outlet blocks AI, skip it” rule misses that split entirely, and misses it in exactly the outlets, national news and major trade press, that make up most tech PR media lists to begin with.
What Should a Tech PR Team Check Before Pitching a Publication?
News publishers block at least one AI crawler in robots.txt at more than five times the rate of the general web, 56.4% versus 10.3%, per HasData. That means most of a typical technology PR target list, which skews toward exactly this kind of outlet, will show a block on paper. The data says that shouldn't move the outlet up or down the list on its own.
| Signal | Predicts citation? | Why |
|---|---|---|
Outlet blocks GPTBot in robots.txt | No | Blocking a training crawler doesn't reliably block the separate bots that power AI citations |
Byline is independent, credible journalism | Yes | Matches the earned-media mechanism behind the large majority of AI citations |
Story was published recently | Yes | Citation odds drop sharply within months of publication |
Outlet's overall domain authority or traffic | Weak on its own | Smaller trade outlets often out-cite bigger generalist sites on their specialty topic |
Outlet allows AI “answer” bots (e.g. OAI-SearchBot, PerplexityBot) | Partially | The split between blocking training bots and allowing answer bots predicts more than a blanket policy does |
The pattern holding across nearly every row: the old media-relations fundamentals, independence, recency, specialization, still decide the outcome. The robots.txt question is mostly noise a PR team can stop checking.
How Does 5W Build a Technology PR Media List for the AI-Search Era?
5WPR's Technology PR practice builds media targets from what a client's category is actually citing across ChatGPT, Perplexity, and Gemini, not from a crawler-policy checklist. That means running a AI Citation Source Audit before the pitch list gets built, so the outlets chosen are the ones already winning the citation in that specific category, whatever their robots.txt happens to say.
5WPR's GEO practice runs that audit on a recurring basis, since which outlets are winning a given category's citations shifts as coverage ages out and new stories enter the mix, the same freshness dynamic that governs every other category these models cite from.
CONCLUSION
A technology PR team crossing outlets off its media list because they block GPTBot is discarding some of the exact publications most likely to show up in an AI answer anyway. 5WPR's Technology PR practice builds media strategy around what's actually being cited, not around a robots.txt flag.





