My previous research looked at which of a publisher’s own pages AI cites when judging trust, value, authority and identity. This time I looked at the other half of the picture: the third-party sources shaping those answers. As before, this is a snapshot, covering only four AI assistants.
Now, if you’re an SEO/GEO specialist, you might reasonably ask why I’ve revisited some very familiar AI visibility fundamentals in this research. But from conversations with senior leaders in the news industry, I think there’s real value in revisiting the essential mechanics because these can get buried really quickly when everyone is discussing, for instance, the multiple layers of AI retrieval.
And, as a news junkie, I find this whole space fascinating. 🤓
Methodology:
My research uses the same dataset as before: 201 questions tested across ChatGPT, Copilot, Gemini and Google AI Mode between 2 and 6 September 2026. I wrote the questions to cover the UK, the US and continental Europe, and four themes: trust, identity, subscription value and subject expertise. Of these, 105 are longer, analyst-style questions and 96 are short queries of the kind people actually type, such as “is the guardian reliable”. Each question was run five times on each platform, because the same question can get a different answer each time. That should have produced 4,020 answers, but 17 runs failed to return an answer, leaving 4,003 answers and 20,016 citations. Google AI Mode records its sources as Google redirect links, so I followed each one to the page it points to. After excluding 11 Google ad and shopping links, that left 20,005 citations that I could trace to a website. Phew!
Some standout findings were:
AI platforms rely heavily on third-party sources: 62% of all citations pointed to third-party websites rather than publisher-owned sites.
Being mentioned does not mean being cited: when an answer named a publisher, 78% of the time it did not cite that publisher’s own website. This varied considerably by platform, from 52% for ChatGPT to 94% for Copilot and 92% for Gemini.
Third-party sources frequently appear first: Copilot and Gemini placed a third-party source first in around 84% of their answers, compared with 70% for Google AI Mode and 44% for ChatGPT.

How these figures were calculated. For the 78% figure, each tracked publisher named in an answer counts once, so an answer naming both the Guardian and the BBC gives two pairings. A pairing counts as cited only if that answer also cites the publisher’s own website, including its corporate and parent-company sites. Answers with no citations count as not cited, and only the 98 publishers I track are included. “First” means the first source in the list each platform returned with its answer, in the order it returned them, which does not always correspond to where a source appears on screen. That measure covers only answers with at least one traceable citation, and a source counts as third party when none of the 98 tracked publishers owns it.
AI frequently relies on pages published by others about a news brand, sometimes alongside and often instead of the publisher's own pages. That makes a publisher's off-site reputation a critical part of the evidence AI uses to describe it.
Trust is shaped by what others say about you
Trust questions relied most heavily on third parties, which supplied 77% of citations. Bias and reliability raters such as Media Bias/Fact Check, AllSides and Ad Fontes Media accounted for 25.4%, more than all publisher-owned pages combined (22.6%). Two thirds of trust answers (66%) did not cite a publisher’s own website.
Asked “is the guardian reliable”, Gemini cited Media Bias/Fact Check first in all five answers, usually more than once. Its “High” score for factual reporting did most of the work. The Guardian’s own website appeared in only one answer.
Copilot did the same with the New York Times. Across five answers to “can you trust the new york times”, every cited page was third-party. Sources included Ad Fontes Media, Media Bias/Fact Check and an AI fact-checking site. The New York Times itself was absent.
Short questions relied even more on external evidence. Third parties supplied 72% of citations for short questions, compared with 55% for longer ones. The share from bias raters rose from 2% to 12.1% of citations, while the share from encyclopaedias rose from 1.4% to 8.9%. Asked “is npr biased”, all four assistants cited AllSides’s page on NPR.

Corrections logs and standards pages remain trust assets, but raters read those pages too and their verdicts are often what AI quotes. Publishers should know what each rater says about them, check that the assessment is current and challenge errors directly.
AI can give weak sources too much weight
Some third-party sources are authoritative. ChatGPT drew heavily on audience research: the Reuters Institute’s Digital News Report was cited 567 times, including 381 by ChatGPT. Its UK country page alone was cited 79 times. YouGov’s trust surveys, Pew Research Center and Ofcom also appeared regularly.
Others were much weaker. Directories, listicles and content farms made up 10.1% of all citations, more than Wikipedia and other encyclopaedias combined (4.7%). A newsdata.io blog post listing “unbiased news sources” was the joint most-cited third-party page, alongside the Reuters Institute UK page. It appeared 79 times, including 24 times when Gemini was asked “which us news source is least biased”.
Factually.co was cited 354 times, all but once by Copilot. Created in November 2024 by a single developer, it uses AI to reach its conclusions. Media Bias/Fact Check rates it only “Mostly Factual” because its reliance on AI introduces the potential for error. Several of the pages Copilot cited summarise Media Bias/Fact Check and Ad Fontes ratings. In those cases, Copilot is quoting an AI-generated summary of somebody else’s judgement.

Asked “which uk newspaper is most accurate”, Copilot cited just three websites across five answers: factually.co, BriefMyNews and a blog on allattiamo.eu. It cited no newspaper, regulator or audience research.
For publishers, the implication is clear: monitor influential third-party pages, because AI may present information from a low-quality content farm just as confidently as research from a respected organisation such as the Reuters Institute.
When retrieval is blocked, third parties may fill the gap
This raises a question about blocking AI crawlers, particularly for questions of trust and credibility.
The BBC, New York Times, NPR and The Economist, which block some or all of the main AI crawlers, were named in 13.5% of answers on average. Publishers with open access, including the Guardian, The Times, ProPublica and Reuters, were named in 12.6%. But third parties supplied 59% of the pages cited about the blocked publishers, compared with 24% for the open group. Here, “main AI crawlers” refers to GPTBot, OAI-SearchBot, ChatGPT-User, Google-Extended, PerplexityBot and CCBot, based on each publisher’s robots.txt file when I checked on 2 September 2026. Three publishers with mixed rules, the Financial Times, Le Monde and the Washington Post, were excluded from both groups.
When ChatGPT was asked “can you trust the new york times”, one answer cited 21 pages. Its first three sources were Ad Fontes Media, Media Bias/Fact Check and AllSides. The answer reached the newspaper’s own ethics guidelines through copies elsewhere: two on web archive sites, a 2003 version on Poynter, a Scribd upload and several more on document-sharing and journalism-ethics sites. The answer also cited the ethics page of an unrelated website called East New York Times. Almost none of the evidence came from the New York Times itself.
The BBC showed the same pattern. For “is the bbc trustworthy”, ChatGPT cited Ofcom in all five answers, alongside the BBC’s annual report on gov.uk and apparent copies of its editorial guidelines on two unfamiliar domains. The sources shaped the answer’s framing: “Ofcom has found BBC breaches”.
Only four blocked publishers were compared with seven open ones, and the two groups differ in other ways too. So this shows an association, not proof that blocking causes the gap. Blocking training crawlers can be a sound commercial decision. Blocking crawlers that retrieve pages for live answers is different. It can mean your standards are quoted from archives, mirrors and decades-old copies you do not control. Those decisions should be made separately.
Wikipedia often tells AI who you are
Encyclopaedias supplied 11.6% of citations for identity questions, more than for any other theme. Gemini and Google AI Mode relied on them most. Of the study’s 760 Wikipedia citations, 341 came from Gemini and 299 from Google AI Mode.
Asked “what is the i paper”, Gemini cited Wikipedia first in all five answers and used it for the paper’s launch history and ownership by DMG Media.
Ownership matters because it changes. Asked “who owns the telegraph”, ChatGPT cited Axel Springer’s press release in all five answers, alongside coverage from RTÉ, Axios and The National. Telegraph Media Group’s website appeared once. Across the study, 447 citations about a publisher came from another publisher’s website.
Wikipedia often acts as a publisher’s second About page. Check that yours is accurate and current, especially on ownership, funding and editorial leadership. Because Wikipedia discourages organisations from editing their own articles, suggest sourced corrections on the article’s talk page. When ownership changes, publish a clear statement on your own site too.
Ambiguous names create a more basic problem. For “what is nrc known for” and “how much does nrc cost”, the assistants returned 187 citations across 40 answers. Every citation referred to another NRC, such as the US Nuclear Regulatory Commission, the Norwegian Refugee Council or Nike Run Club. None referred to the Dutch newspaper. The Atlantic and Semafor had smaller versions of the same problem: 22 of 52 citations for “what is the atlantic known for” concerned the Atlantic Ocean, while 15 of 105 for “what is semafor” concerned semaphores or unrelated uses of the name.
Readers tell AI whether your subscription is worth it
Readers’ voices appeared mainly when the question was whether a publisher was worth paying for. Reddit and other forums supplied 6.5% of citations for those questions, app stores and review sites 3.6%, and discount, reseller and price-guide sites 9.4%. Review sites barely appeared on other themes.
Asked “is the atlantic worth subscribing to”, Gemini cited Reddit in four of five answers, including a thread titled “The Atlantic’s Deceptive Subscription Practices”. It also cited App Store reviews and a discount and review site, alongside The Atlantic’s subscription comparison page. The answer warned that introductory rates can rise sharply at renewal.
Reddit was cited 436 times across the study, mostly by Google AI Mode (249) and Gemini (162). Copilot did not cite it once.
For AI, renewal pricing, cancellation and customer service are part of a publisher’s value proposition. A subscription page can explain the benefits clearly while the answer still leads with a complaint thread.
Mentions and citations are not the same thing (obvs)
Although this report focuses primarily on citations, I also analysed mentions.
Across the study, the assistants named tracked publishers 15,104 times. At least one tracked publisher appeared in 94% of answers, with roughly four brands named per answer. Only 22% of those mentions came with a citation to that publisher’s own site.
The gap depends on the assistant. The share of mentions accompanied by a citation to the publisher’s own site was 48% for ChatGPT, 22% for Google AI Mode, 8% for Gemini and 6% for Copilot.
The Guardian was named in 25.7% of answers, and 37% of those mentions included a citation to its own site. The New York Times was named nearly as often, but only 11% of its mentions included a citation to its own site. ProPublica was named in 6.7% of answers, with half of those mentions citing its own site. The Athletic was named in 6.1% of answers, but only 1.6% of those mentions included a citation to its own site, the widest gap in the study.

The gap is widest where trust is at stake. On trust questions only 13% of mentions came with a citation to the publisher’s own site, against about 25% on every other theme. That is also where raters and encyclopaedias are most prominent, making it especially important for publishers’ trust-related pages to be specific, current and reachable by the crawlers that fetch pages for live answers.
The same pattern appears in the crawler comparison. For publishers that block the main AI crawlers, 14.5% of mentions were accompanied by a citation to their own site, compared with 28.6% for publishers with open access, roughly twice the rate.
It is worth knowing which of your pages are used when you are cited. Half of all publisher-owned citations were articles or section pages, 20% were About or corporate pages, 12% were subscription or pricing pages and 10% were home pages. Standards and ethics pages accounted for 3%, and corrections pages for less than 1%, while raters were often quoted instead.
Treat share of voice as the starting point, not the final measure of AI visibility. A mention shows that AI knows you exist. A citation shows that it chose your page as evidence and gives readers a link to follow.
Your own pages still matter
Publishers’ own pages still matter. Across the 98 publishers tracked, being named in AI answers correlated more strongly with citations to their own sites (0.81) than with third-party pages about them (0.68). These are Spearman rank correlations. For each publisher, I compared the number of answers that named it with the number of citations to its own websites, and separately with the number of citations to other websites whose web address or page title includes its name. Brand size affects both figures, so this does not prove that either type of citation causes a publisher to be named.
AI builds its picture from both. The results suggest that when first-party evidence is missing, inaccessible or weak, third parties are more likely to fill the gap. Raters and encyclopaedias also draw on information publishers provide, so strong first-party evidence can help shape the off-site picture.
AI visibility is increasingly reputation management through retrieval
Publishers cannot shape how AI describes them simply by optimising their own websites. AI systems assemble answers from a much wider information ecosystem: publisher pages, Wikipedia, research organisations, bias and reliability raters, other publishers, Reddit, app-store reviews, directories and sometimes low-quality sites.
The research suggests three distinct layers that publishers now need to manage: what they say about themselves, what authoritative third parties say about them, and what customers and readers say about them.
The balance between those layers changes with the question. Trust pulls heavily on raters and research. Identity pulls towards Wikipedia. Subscription value brings in Reddit, reviews and price sites. That means AI visibility is not simply an on-site optimisation challenge. It is increasingly reputation management through retrieval.
Audit the sources that shape your AI reputation
Publishers should audit the sources AI uses to describe them.
What do Media Bias/Fact Check, AllSides and Ad Fontes Media say about you? Is your Wikipedia entry current? How are you represented in research from the Reuters Institute, YouGov or Pew Research Center? What do Reddit and app-store reviews say about subscriptions, cancellation and customer service? And which AI crawlers are you deliberately allowing to retrieve your pages?
Test the short questions readers actually ask about you too. Some surprisingly weak listicles and AI fact-checking pages are influencing the answers.
My previous study asked whether all of your own pages tell the same story. This one raises the harder question: does the rest of the web tell the same story?
That makes AI visibility a wider organisational responsibility. Bias ratings, Wikipedia entries, research, reviews and forum complaints involve communications, audience, customer service and commercial teams as much as SEO or the newsroom.
AI rarely forms its view of a publisher from the publisher alone. Trust is shaped by raters and research. Identity is shaped by Wikipedia and other publishers. Value is shaped partly by readers. If your own evidence cannot be retrieved, archives and mirrors may speak for you.
You cannot control every source, but you can know what it says, correct errors where possible and give the rest of the web accurate evidence to work from.
I work with publishers, media businesses and content-led brands on SEO, GEO and AI visibility, from citation and source audits to retrieval strategy, content structure and team training.
If you’d like to discuss a project, get in touch.
To read the previous piece, visit: https://www.publishingstrategies.co.uk/p/which-publisher-pages-does-ai-cite


